Skip to content

Portal | Level: L1: Foundations | Topics: Helm | Domain: DevOps & Tooling

Lab Runtime 05 — Helm Upgrade & Rollback

Objective

Introduce a bad Helm values change, upgrade the release, recover via rollback, and verify the application is healthy again.

Prerequisites

  • make deploy-all has been run and the grokdevops app is deployed and healthy
  • Helm 3 installed
  • kubectl configured for the target cluster

Steps

  1. Break: Run ./break.sh to create a bad values override with a nonexistent image tag and perform a Helm upgrade. The upgrade will fail or leave pods in ImagePullBackOff.
  2. Observe: Check pod status with kubectl get pods -n grokdevops and Helm history with helm history grokdevops -n grokdevops.
  3. Fix: Run ./fix.sh to rollback to the previous working Helm revision.
  4. Verify: Run ./verify.sh to confirm the deployment is healthy with ready replicas.
  5. Teardown: Run ./teardown.sh to ensure clean Helm state and remove temporary files.

Expected Observations

  • kubectl get pods -n grokdevops shows new pods stuck in ImagePullBackOff or ErrImagePull because the image tag does not exist in the registry.
  • kubectl describe pod <pod-name> -n grokdevops shows events like Failed to pull image and image not found.
  • helm history grokdevops -n grokdevops shows the latest revision with status deployed (Helm marks it deployed even though pods are failing) or failed if the upgrade timed out.

Wrong Turns

  1. Deleting the failing pods with kubectl delete pod — The deployment controller recreates pods from the same spec, which still references the nonexistent image tag. New pods fail identically.
  2. Editing the deployment directly with kubectl edit — This fixes the running state but creates Helm state drift. The next helm upgrade may revert your fix or conflict with stored release state.
  3. Deleting the Helm release entirely — This tears down the whole application unnecessarily. helm rollback is the correct, surgical recovery that preserves release history.

Minimal Explanation

Helm stores each upgrade as a numbered revision in a Kubernetes Secret. When you run helm upgrade with a bad image tag, Helm renders the new templates, applies them to the cluster, and creates a new revision. The deployment controller then tries to roll out pods using the new image. The kubelet on each node attempts to pull the image from the registry. When the tag does not exist, the pull fails and the pod enters ImagePullBackOff with exponential backoff. helm rollback re-applies the templates from a previous good revision, restoring the working image tag and triggering a new rollout.

Transfer Pattern

  • Fat-finger image tag in CI/CD: A typo in the image tag variable (e.g., v1.2.3 vs v1.23) causes a registry lookup failure across all new pods.
  • Registry outage during deploy: If the container registry is temporarily unavailable, image pulls fail the same way. Rollback restores the previous revision while you fix the registry issue.

See Also

  • training/library/runbooks/cicd/helm_upgrade_failed.md
  • training/interview-scenarios/05-helm-upgrade-broke-prod.md

Solution (spoilers)

See training/library/solutions/labs/lab-runtime-05.md for hints and explanation.

Teardown

./teardown.sh

Or reset the entire environment:

make undeploy-all

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)