Portal | Level: L1: Foundations | Topics: OOMKilled, Kubernetes Core | Domain: Kubernetes
Lab Runtime 08 — Resource Limits and OOMKilled¶
Objective¶
Set extremely low memory limits on the grokdevops deployment to trigger OOMKilled, observe the failure, then fix with proper resource limits.
Prerequisites¶
make deploy-allhas been run and the grokdevops app is deployed and healthy- Helm 3 installed
- kubectl configured for the target cluster
Steps¶
- Break: Run
./break.shto patch the deployment with an extremely low memory limit (4Mi). The container will be OOMKilled almost immediately. - Observe: Watch pods restart with
kubectl get pods -n grokdevops -w. Usekubectl describe podto see the OOMKilled reason in the Last State. - Fix: Run
./fix.shto restore reasonable memory limits (128Mi request, 256Mi limit). - Verify: Run
./verify.shto confirm pods are running without OOMKills. - Teardown: Run
./teardown.shto restore the deployment to its Helm-declared state.
Expected Observations¶
kubectl get pods -n grokdevopsshows pods inCrashLoopBackOffwith a high restart count.kubectl describe pod <pod-name> -n grokdevopsshowsLast State: TerminatedwithReason: OOMKilledandExit Code: 137.- Pod events show the container starting and being killed repeatedly, with increasing backoff intervals between restarts.
Wrong Turns¶
- Blaming an application crash or bug — Exit code 137 means the process received SIGKILL from the kernel OOM killer. This is an infrastructure-level kill, not an application exception or unhandled error.
- Removing memory limits entirely — This allows the container to consume unbounded memory, potentially starving other pods on the same node and causing node-level instability.
- Only increasing memory requests without raising limits — Requests affect scheduling (which node the pod lands on), but the limit is what the kernel enforces. If the limit stays at 4Mi, the container is still OOM-killed regardless of the request value.
Minimal Explanation¶
When a pod spec includes a memory limit, Kubernetes configures a Linux cgroup with that limit. The cgroup constrains the total memory (RSS + cache) the container's processes can use. When the container's memory consumption exceeds the cgroup limit, the kernel's OOM killer activates and sends SIGKILL (signal 9) to the primary process. The container exits with code 137 (128 + 9). Kubernetes sees the terminated container and restarts it per the pod's restart policy, but if the container immediately exceeds the limit again, it enters CrashLoopBackOff with exponentially increasing delays between restart attempts.
Transfer Pattern¶
- Memory-intensive batch jobs: Data processing or report generation tasks that load large datasets into memory can exceed limits set for normal request-serving workloads.
- Java heap sizing: A JVM configured with
-Xmxlarger than the container's memory limit will be OOM-killed when heap usage approaches max. The JVM heap must fit within the container limit with room for off-heap memory. - Python data processing: Libraries like pandas or numpy can allocate large arrays that push memory usage above container limits, especially when processing files of variable size.
See Also¶
training/library/runbooks/kubernetes/oom-kill.mdtraining/interview-scenarios/08-pods-oomkilled.md
Solution (spoilers)¶
See training/library/solutions/labs/lab-runtime-08.md for hints and explanation.
Teardown¶
Or reset the entire environment:
Wiki Navigation¶
Prerequisites¶
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)
Related Content¶
- Case Study: Node Pressure Evictions (Case Study, L2) — Kubernetes Core, OOMKilled
- Case Study: Pod OOMKilled — Memory Leak in Sidecar, Fix Is Helm Values (Case Study, L2) — Kubernetes Core, OOMKilled
- Ops Archaeology: The Session Store That Keeps Dying (Case Study, L2) — Kubernetes Core, OOMKilled
- Runbook: OOMKilled Container (Runbook, L1) — Kubernetes Core, OOMKilled
- Adversarial Interview Gauntlet (30 sequences) (Scenario, L2) — Kubernetes Core
- Case Study: Alert Storm — Flapping Health Checks (Case Study, L2) — Kubernetes Core
- Case Study: Canary Deploy Routing to Wrong Backend — Ingress Misconfigured (Case Study, L2) — Kubernetes Core
- Case Study: CrashLoopBackOff No Logs (Case Study, L1) — Kubernetes Core
- Case Study: DNS Looks Broken — TLS Expired, Fix Is Cert-Manager (Case Study, L2) — Kubernetes Core
- Case Study: DaemonSet Blocks Eviction (Case Study, L2) — Kubernetes Core