Give the init container resource requests and limits - #1259
gregorjerse wants to merge 1 commit into
Conversation
ef5b384 to
5c76db0
Compare
The init container had none, so Kubernetes ran it with the minimal CPU share and the highest OOM score in the pod: the download of the inputs was starved on a busy node and killed first under memory pressure. It now requests 1 CPU and 1Gi and is limited to 2 CPUs and 2Gi, enough for the three download workers with their in-flight parts. The settings FLOW_KUBERNETES_INIT_CONTAINER_LIMITS and FLOW_KUBERNETES_INIT_CONTAINER_REQUESTS override the defaults.
5c76db0 to
62a90d0
Compare
dblenkus
left a comment
There was a problem hiding this comment.
Right problem — an init container with no requests really does get cpu.weight 1 and the worst oom_score_adj in the pod. Two things about the numbers.
-
"these values never raise it" does not hold for the default process.
cores=1, memory=4096with the default overcommit{"cpu": 0.8, "memory": 0.8}gives 0.8 CPU for processing plus 0.1 for the communicator = 0.9. The init container's 1 CPU raises the pod's effective CPU request for every single-core job, and by more whereverFLOW_KUBERNETES_OVERCOMMITlowers the factor for a scheduling class. Memory is fine (3276 MiB + 256 M ≫ 1 Gi). Requesting 0.5 CPU would keep the claim true. -
The 2 Gi limit is not derived from the knobs it depends on. Downloader memory ≈
GENESIS_MAX_DOWNLOAD_THREADS × 10(botomax_concurrency;use_threads = Trueis hardcoded inS3Connector)× chunk_size, andchunk_size = max(8 MiB, size/10000). The defaults give the ~240 MiB in the docstring, but an input above ~84 GB pushes chunk_size past 8 MiB (200 GB → 600 MiB, 500 GB → 1.5 GiB), and raising the thread count multiplies it. Going from no limit to 2 Gi converts "slow because starved" into OOMKilled, whichbackoffLimit: 0will not retry. Worth scaling the limit from the thread count, or at least stating the assumption. -
_init_container_resourcesdoes not useself.
Note this and #1260 both open an Unreleased/Changed section — CHANGELOG conflict for whichever lands second.
Independent of #1256, #1257, #1258 and #1260; branches off master.
Problem
The init container had no resource requests or limits. Kubernetes then gives it the minimal CPU share (
cpu.shares2,cpu.weight1), so on a busy node the download of the inputs is starved by every pod that has a request, and itsoom_score_adjis the highest in the pod, so it is the first process the kernel kills under memory pressure.Fix
The init container requests 1 CPU and 1Gi of memory, with limits of 2 CPUs and 2Gi. The download runs a few Python threads and needs far less than the processing container: the memory covers the three download workers with their ten in-flight 8 MiB parts each plus the s3transfer io queues. The settings
FLOW_KUBERNETES_INIT_CONTAINER_LIMITSandFLOW_KUBERNETES_INIT_CONTAINER_REQUESTSoverride the defaults, in the shape of the communicator settings. The pod's effective request is the maximum of the init containers and the sum of the app containers, so these values never raise it.Verification
resolwe.flow.tests.test_utils.KubernetesTestCasepasses; the new test fails against master.