Skip to content

Portal | Level: L2: Operations | Topics: Incident Response, Kubernetes Core | Domain: Kubernetes

Chaos Scripts

Safe, reversible chaos experiments for practicing incident response on a live cluster.

Rules

  1. All scripts are namespace-scoped (default: grokdevops)
  2. All scripts support --dry-run (preview) and --yes (confirm destructive action)
  3. All faults are reversible, but restoration may be automatic or require the cleanup command printed by the script
  4. Scripts never operate outside approved namespaces (grokdevops, monitoring, kube-system) unless --i-know-what-im-doing is passed

Prerequisites

  • Cluster running with make deploy-all completed
  • kubectl configured and pointing at the cluster

Available Scripts

Script What it does Default namespace
kill_pods.sh Delete pods by label (they restart) grokdevops
break_readiness.sh Patch deployment to break /health probe grokdevops
cpu_stress.sh Run a CPU stress pod grokdevops
mem_stress.sh Run a memory stress pod (may trigger OOM) grokdevops
toggle_networkpolicy.sh Apply/remove a restrictive NetworkPolicy grokdevops
scale_to_zero.sh Scale deployment to 0, then restore grokdevops
inject_bad_configmap.sh Inject a ConfigMap; print manual restore commands grokdevops

Usage

# Preview what would happen
./training/interactive/chaos/scripts/kill_pods.sh --dry-run

# Execute (requires --yes)
./training/interactive/chaos/scripts/kill_pods.sh --yes

# Different namespace
./training/interactive/chaos/scripts/kill_pods.sh --yes --namespace monitoring

# Via Makefile
make chaos LIST=1

Safety

  • Without --yes, scripts only print what they would do
  • Approved namespaces: grokdevops, monitoring, argocd, kube-system
  • To use on other namespaces: pass --i-know-what-im-doing
  • kill_pods.sh relies on the workload controller to recreate deleted pods; scale_to_zero.sh restores the original replica count. The CPU and memory stress processes end after their configured duration, but their completed pods can be removed with the printed cleanup command.
  • break_readiness.sh, inject_bad_configmap.sh, and toggle_networkpolicy.sh apply leave the fault active and print the command needed to restore it. Run that cleanup before ending the exercise.

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)