Skip to content

Portal | Level: L1: Foundations | Topics: HPA / Autoscaling, Kubernetes Core | Domain: Kubernetes

Lab Runtime 02 — HPA Live Scaling

Objective

Generate load against the grokdevops app, confirm the Horizontal Pod Autoscaler (HPA) reacts by scaling up replicas, then stop load and watch it scale back down. Validate scaling via kubectl top and HPA status.

Prerequisites

  • Cluster running with the grokdevops app deployed: make deploy-all
  • metrics-server installed and operational (required for HPA)
  • kubectl configured for the target cluster

Steps

  1. Generate load — Run ./break.sh. This ensures an HPA exists and starts a busybox pod that hammers the /health endpoint in a tight loop.

  2. Observe scale-up — Watch the HPA react (metrics may take ~60s to propagate):

    kubectl get hpa -n grokdevops -w
    kubectl top pods -n grokdevops
    

  3. Verify — Run ./verify.sh to check current replica count.

  4. Stop load — Run ./fix.sh to delete the load generator pod. The HPA will scale down after its cooldown period (~5 minutes).

Expected Observations

  • kubectl get hpa -n grokdevops shows the CPU utilization percentage climbing above the target threshold.
  • The HPA's REPLICAS column increases from the initial count as CPU load rises.
  • kubectl top pods -n grokdevops shows elevated CPU usage on the load-receiving pods.
  • After the load generator is removed, CPU drops and replicas gradually decrease after the HPA cooldown period (~5 minutes).

Wrong Turns

  1. Manually scaling with kubectl scale — The HPA continuously reconciles replica count to its computed target, so manual scaling is overridden within seconds.
  2. Removing the HPA to control scaling yourself — This eliminates autoscaling entirely, defeating the purpose. In production, you lose the ability to react to traffic changes automatically.
  3. Setting a CPU target percentage that is too low or too high — A target too low (e.g., 5%) causes excessive scaling and resource waste; a target too high (e.g., 95%) means the HPA barely reacts before pods are saturated.

Minimal Explanation

The HPA controller queries metrics-server every 15 seconds for pod CPU usage. It computes the desired replica count as: desiredReplicas = ceil(currentReplicas * (currentCPUUtilization / targetCPUUtilization)) For this formula to work, pods must have CPU resource requests defined (the HPA computes utilization as a percentage of the request). Without metrics-server or without resource requests, the HPA reports <unknown> for CPU and cannot scale. The controller also enforces a stabilization window to avoid thrashing — scale-up is fast, but scale-down waits ~5 minutes by default.

Transfer Pattern

  • Black Friday traffic spikes: Sudden surges in web traffic cause CPU to spike; the HPA scales pods up to absorb load and scales back down when the event ends.
  • Batch processing jobs: CPU-intensive workloads like report generation or data imports can trigger horizontal scaling to maintain responsiveness for other requests.

See Also

  • training/library/runbooks/kubernetes/hpa_not_scaling.md
  • training/interview-scenarios/02-hpa-not-scaling.md

Solution (spoilers)

See training/library/solutions/labs/lab-runtime-02.md for hints and explanation.

Teardown

./teardown.sh

Or reset the entire environment:

make undeploy-all

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)