Portal | Level: L1: Foundations | Topics: Probes (Liveness/Readiness), Kubernetes Core | Domain: Kubernetes
Lab Runtime 01 — Rollout Probe Failure¶
Objective¶
Break the readiness probe via a Helm values override (or direct patch), watch the rollout stall because new pods never become ready, then fix it by restoring the correct probe path.
Prerequisites¶
- Cluster running with the grokdevops app deployed:
make deploy-all kubectlconfigured for the target cluster
Steps¶
-
Break it — Run
./break.sh. This patches the deployment's readiness probe to hit/nonexistent, which always returns a non-200 response. New pods will never pass readiness checks and the rollout will stall. -
Observe — Watch the rollout hang:
You should see new pods stuck inkubectl rollout status deployment/grokdevops -n grokdevops --timeout=30s kubectl get pods -n grokdevops0/1 Running(not READY). -
Fix it — Run
./fix.sh. This restores the readiness probe path to/healthand waits for the rollout to complete. -
Verify — Run
./verify.shto confirm all replicas are ready.
Expected Observations¶
kubectl get pods -n grokdevopsshows new pods in0/1 Running— the container is running but the readiness gate has not passed.kubectl describe pod <pod-name> -n grokdevopsshows the readiness probe configuration pointing at/nonexistentand a conditionReady: False.- Events on the pod include repeated
Readiness probe failed: HTTP probe failed with statuscode: 404messages. kubectl rollout status deployment/grokdevops -n grokdevopsreports the rollout is waiting and never completes.
Wrong Turns¶
- Deleting the failing pods — Kubernetes will respawn new pods from the same deployment spec, so the new pods have the same broken probe path and fail identically.
- Scaling the deployment to zero and back — This drains and recreates pods, but the deployment spec still carries the wrong probe path, so the new pods fail the same way.
- Restarting the deployment with
kubectl rollout restart— This triggers a new rollout with the same misconfigured probe, producing a fresh set of pods that also never become ready.
Minimal Explanation¶
Kubernetes readiness probes gate whether a pod is added to Service endpoints.
The kubelet periodically sends an HTTP request to the configured probe path.
When the probe target returns a non-2xx status (or times out), the pod's
Ready condition stays False. The Endpoints controller sees this and
removes the pod from the Service's endpoint list, so no traffic is routed
to it. During a rolling update, the deployment controller waits for new
pods to become ready before scaling down old ones. If the new pods never
pass readiness, the rollout stalls indefinitely, leaving the old pods
serving traffic.
Transfer Pattern¶
- Bad health-check endpoint after code deploy: A developer renames or removes the
/healthroute but forgets to update the probe path in the Helm values or deployment manifest. - Slow-starting applications: Apps that take longer than the probe's
initialDelaySecondsto boot will fail readiness checks on startup, causing rolling updates to stall until the timing is corrected.
See Also¶
training/library/runbooks/kubernetes/readiness_probe_failed.mdtraining/interview-scenarios/01-deployment-stuck-progressing.md
Solution (spoilers)¶
See training/library/solutions/labs/lab-runtime-01.md for hints and explanation.
Teardown¶
Or reset the entire environment:
Wiki Navigation¶
Prerequisites¶
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)
Related Content¶
- Interview: Deployment Stuck Progressing (Scenario, L2) — Kubernetes Core, Probes (Liveness/Readiness)
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1) — Kubernetes Core, Probes (Liveness/Readiness)
- Runbook: Readiness Probe Failed (Runbook, L1) — Kubernetes Core, Probes (Liveness/Readiness)
- Skillcheck: Kubernetes (Assessment, L1) — Kubernetes Core, Probes (Liveness/Readiness)
- Track: Kubernetes Core (Reference, L1) — Kubernetes Core, Probes (Liveness/Readiness)
- Adversarial Interview Gauntlet (30 sequences) (Scenario, L2) — Kubernetes Core
- Case Study: Alert Storm — Flapping Health Checks (Case Study, L2) — Kubernetes Core
- Case Study: Canary Deploy Routing to Wrong Backend — Ingress Misconfigured (Case Study, L2) — Kubernetes Core
- Case Study: CrashLoopBackOff No Logs (Case Study, L1) — Kubernetes Core
- Case Study: DNS Looks Broken — TLS Expired, Fix Is Cert-Manager (Case Study, L2) — Kubernetes Core