Portal | Level: L1: Foundations | Topics: Helm | Domain: DevOps & Tooling
Lab Runtime 05 — Helm Upgrade & Rollback¶
Objective¶
Introduce a bad Helm values change, upgrade the release, recover via rollback, and verify the application is healthy again.
Prerequisites¶
make deploy-allhas been run and the grokdevops app is deployed and healthy- Helm 3 installed
- kubectl configured for the target cluster
Steps¶
- Break: Run
./break.shto create a bad values override with a nonexistent image tag and perform a Helm upgrade. The upgrade will fail or leave pods inImagePullBackOff. - Observe: Check pod status with
kubectl get pods -n grokdevopsand Helm history withhelm history grokdevops -n grokdevops. - Fix: Run
./fix.shto rollback to the previous working Helm revision. - Verify: Run
./verify.shto confirm the deployment is healthy with ready replicas. - Teardown: Run
./teardown.shto ensure clean Helm state and remove temporary files.
Expected Observations¶
kubectl get pods -n grokdevopsshows new pods stuck inImagePullBackOfforErrImagePullbecause the image tag does not exist in the registry.kubectl describe pod <pod-name> -n grokdevopsshows events likeFailed to pull imageandimage not found.helm history grokdevops -n grokdevopsshows the latest revision with statusdeployed(Helm marks it deployed even though pods are failing) orfailedif the upgrade timed out.
Wrong Turns¶
- Deleting the failing pods with
kubectl delete pod— The deployment controller recreates pods from the same spec, which still references the nonexistent image tag. New pods fail identically. - Editing the deployment directly with
kubectl edit— This fixes the running state but creates Helm state drift. The nexthelm upgrademay revert your fix or conflict with stored release state. - Deleting the Helm release entirely — This tears down the whole application unnecessarily.
helm rollbackis the correct, surgical recovery that preserves release history.
Minimal Explanation¶
Helm stores each upgrade as a numbered revision in a Kubernetes Secret.
When you run helm upgrade with a bad image tag, Helm renders the new
templates, applies them to the cluster, and creates a new revision. The
deployment controller then tries to roll out pods using the new image.
The kubelet on each node attempts to pull the image from the registry.
When the tag does not exist, the pull fails and the pod enters
ImagePullBackOff with exponential backoff. helm rollback re-applies
the templates from a previous good revision, restoring the working image
tag and triggering a new rollout.
Transfer Pattern¶
- Fat-finger image tag in CI/CD: A typo in the image tag variable (e.g.,
v1.2.3vsv1.23) causes a registry lookup failure across all new pods. - Registry outage during deploy: If the container registry is temporarily unavailable, image pulls fail the same way. Rollback restores the previous revision while you fix the registry issue.
See Also¶
training/library/runbooks/cicd/helm_upgrade_failed.mdtraining/interview-scenarios/05-helm-upgrade-broke-prod.md
Solution (spoilers)¶
See training/library/solutions/labs/lab-runtime-05.md for hints and explanation.
Teardown¶
Or reset the entire environment:
Wiki Navigation¶
Prerequisites¶
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)
Related Content¶
- Case Study: Pod OOMKilled — Memory Leak in Sidecar, Fix Is Helm Values (Case Study, L2) — Helm
- Helm (Topic Pack, L1) — Helm
- Helm Drills (Drill, L1) — Helm
- Helm Flashcards (CLI) (flashcard_deck, L1) — Helm
- Incident Simulator (18 scenarios) (CLI) (Exercise Set, L2) — Helm
- Interview: Helm Upgrade Broke Prod (Scenario, L2) — Helm
- Runbook: Helm Upgrade Failed (Runbook, L1) — Helm
- Skillcheck: Helm & Release Ops (Assessment, L1) — Helm
- Track: Helm & Release Ops (Reference, L1) — Helm