Skip to content

Portal | Level: L1: Foundations | Topics: Probes (Liveness/Readiness), Kubernetes Core | Domain: Kubernetes

Lab Runtime 01 — Rollout Probe Failure

Objective

Break the readiness probe via a Helm values override (or direct patch), watch the rollout stall because new pods never become ready, then fix it by restoring the correct probe path.

Prerequisites

  • Cluster running with the grokdevops app deployed: make deploy-all
  • kubectl configured for the target cluster

Steps

  1. Break it — Run ./break.sh. This patches the deployment's readiness probe to hit /nonexistent, which always returns a non-200 response. New pods will never pass readiness checks and the rollout will stall.

  2. Observe — Watch the rollout hang:

    kubectl rollout status deployment/grokdevops -n grokdevops --timeout=30s
    kubectl get pods -n grokdevops
    
    You should see new pods stuck in 0/1 Running (not READY).

  3. Fix it — Run ./fix.sh. This restores the readiness probe path to /health and waits for the rollout to complete.

  4. Verify — Run ./verify.sh to confirm all replicas are ready.

Expected Observations

  • kubectl get pods -n grokdevops shows new pods in 0/1 Running — the container is running but the readiness gate has not passed.
  • kubectl describe pod <pod-name> -n grokdevops shows the readiness probe configuration pointing at /nonexistent and a condition Ready: False.
  • Events on the pod include repeated Readiness probe failed: HTTP probe failed with statuscode: 404 messages.
  • kubectl rollout status deployment/grokdevops -n grokdevops reports the rollout is waiting and never completes.

Wrong Turns

  1. Deleting the failing pods — Kubernetes will respawn new pods from the same deployment spec, so the new pods have the same broken probe path and fail identically.
  2. Scaling the deployment to zero and back — This drains and recreates pods, but the deployment spec still carries the wrong probe path, so the new pods fail the same way.
  3. Restarting the deployment with kubectl rollout restart — This triggers a new rollout with the same misconfigured probe, producing a fresh set of pods that also never become ready.

Minimal Explanation

Kubernetes readiness probes gate whether a pod is added to Service endpoints. The kubelet periodically sends an HTTP request to the configured probe path. When the probe target returns a non-2xx status (or times out), the pod's Ready condition stays False. The Endpoints controller sees this and removes the pod from the Service's endpoint list, so no traffic is routed to it. During a rolling update, the deployment controller waits for new pods to become ready before scaling down old ones. If the new pods never pass readiness, the rollout stalls indefinitely, leaving the old pods serving traffic.

Transfer Pattern

  • Bad health-check endpoint after code deploy: A developer renames or removes the /health route but forgets to update the probe path in the Helm values or deployment manifest.
  • Slow-starting applications: Apps that take longer than the probe's initialDelaySeconds to boot will fail readiness checks on startup, causing rolling updates to stall until the timing is corrected.

See Also

  • training/library/runbooks/kubernetes/readiness_probe_failed.md
  • training/interview-scenarios/01-deployment-stuck-progressing.md

Solution (spoilers)

See training/library/solutions/labs/lab-runtime-01.md for hints and explanation.

Teardown

./teardown.sh

Or reset the entire environment:

make undeploy-all

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)