Skip to content

Portal | Level: L1: Foundations | Topics: OOMKilled, Kubernetes Core | Domain: Kubernetes

Lab Runtime 08 — Resource Limits and OOMKilled

Objective

Set extremely low memory limits on the grokdevops deployment to trigger OOMKilled, observe the failure, then fix with proper resource limits.

Prerequisites

  • make deploy-all has been run and the grokdevops app is deployed and healthy
  • Helm 3 installed
  • kubectl configured for the target cluster

Steps

  1. Break: Run ./break.sh to patch the deployment with an extremely low memory limit (4Mi). The container will be OOMKilled almost immediately.
  2. Observe: Watch pods restart with kubectl get pods -n grokdevops -w. Use kubectl describe pod to see the OOMKilled reason in the Last State.
  3. Fix: Run ./fix.sh to restore reasonable memory limits (128Mi request, 256Mi limit).
  4. Verify: Run ./verify.sh to confirm pods are running without OOMKills.
  5. Teardown: Run ./teardown.sh to restore the deployment to its Helm-declared state.

Expected Observations

  • kubectl get pods -n grokdevops shows pods in CrashLoopBackOff with a high restart count.
  • kubectl describe pod <pod-name> -n grokdevops shows Last State: Terminated with Reason: OOMKilled and Exit Code: 137.
  • Pod events show the container starting and being killed repeatedly, with increasing backoff intervals between restarts.

Wrong Turns

  1. Blaming an application crash or bug — Exit code 137 means the process received SIGKILL from the kernel OOM killer. This is an infrastructure-level kill, not an application exception or unhandled error.
  2. Removing memory limits entirely — This allows the container to consume unbounded memory, potentially starving other pods on the same node and causing node-level instability.
  3. Only increasing memory requests without raising limits — Requests affect scheduling (which node the pod lands on), but the limit is what the kernel enforces. If the limit stays at 4Mi, the container is still OOM-killed regardless of the request value.

Minimal Explanation

When a pod spec includes a memory limit, Kubernetes configures a Linux cgroup with that limit. The cgroup constrains the total memory (RSS + cache) the container's processes can use. When the container's memory consumption exceeds the cgroup limit, the kernel's OOM killer activates and sends SIGKILL (signal 9) to the primary process. The container exits with code 137 (128 + 9). Kubernetes sees the terminated container and restarts it per the pod's restart policy, but if the container immediately exceeds the limit again, it enters CrashLoopBackOff with exponentially increasing delays between restart attempts.

Transfer Pattern

  • Memory-intensive batch jobs: Data processing or report generation tasks that load large datasets into memory can exceed limits set for normal request-serving workloads.
  • Java heap sizing: A JVM configured with -Xmx larger than the container's memory limit will be OOM-killed when heap usage approaches max. The JVM heap must fit within the container limit with room for off-heap memory.
  • Python data processing: Libraries like pandas or numpy can allocate large arrays that push memory usage above container limits, especially when processing files of variable size.

See Also

  • training/library/runbooks/kubernetes/oom-kill.md
  • training/interview-scenarios/08-pods-oomkilled.md

Solution (spoilers)

See training/library/solutions/labs/lab-runtime-08.md for hints and explanation.

Teardown

./teardown.sh

Or reset the entire environment:

make undeploy-all

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)