Portal | Level: L2: Operations | Topics: Incident Response, Kubernetes Core | Domain: Kubernetes
Chaos Scripts¶
Safe, reversible chaos experiments for practicing incident response on a live cluster.
Rules¶
- All scripts are namespace-scoped (default:
grokdevops) - All scripts support
--dry-run(preview) and--yes(confirm destructive action) - All faults are reversible, but restoration may be automatic or require the cleanup command printed by the script
- Scripts never operate outside approved namespaces (
grokdevops,monitoring,kube-system) unless--i-know-what-im-doingis passed
Prerequisites¶
- Cluster running with
make deploy-allcompleted kubectlconfigured and pointing at the cluster
Available Scripts¶
| Script | What it does | Default namespace |
|---|---|---|
kill_pods.sh |
Delete pods by label (they restart) | grokdevops |
break_readiness.sh |
Patch deployment to break /health probe | grokdevops |
cpu_stress.sh |
Run a CPU stress pod | grokdevops |
mem_stress.sh |
Run a memory stress pod (may trigger OOM) | grokdevops |
toggle_networkpolicy.sh |
Apply/remove a restrictive NetworkPolicy | grokdevops |
scale_to_zero.sh |
Scale deployment to 0, then restore | grokdevops |
inject_bad_configmap.sh |
Inject a ConfigMap; print manual restore commands | grokdevops |
Usage¶
# Preview what would happen
./training/interactive/chaos/scripts/kill_pods.sh --dry-run
# Execute (requires --yes)
./training/interactive/chaos/scripts/kill_pods.sh --yes
# Different namespace
./training/interactive/chaos/scripts/kill_pods.sh --yes --namespace monitoring
# Via Makefile
make chaos LIST=1
Safety¶
- Without
--yes, scripts only print what they would do - Approved namespaces:
grokdevops,monitoring,argocd,kube-system - To use on other namespaces: pass
--i-know-what-im-doing kill_pods.shrelies on the workload controller to recreate deleted pods;scale_to_zero.shrestores the original replica count. The CPU and memory stress processes end after their configured duration, but their completed pods can be removed with the printed cleanup command.break_readiness.sh,inject_bad_configmap.sh, andtoggle_networkpolicy.sh applyleave the fault active and print the command needed to restore it. Run that cleanup before ending the exercise.
Wiki Navigation¶
Prerequisites¶
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)
Related Content¶
- Adversarial Interview Gauntlet (30 sequences) (Scenario, L2) — Kubernetes Core
- Case Study: Alert Storm — Flapping Health Checks (Case Study, L2) — Kubernetes Core
- Case Study: Canary Deploy Routing to Wrong Backend — Ingress Misconfigured (Case Study, L2) — Kubernetes Core
- Case Study: CrashLoopBackOff No Logs (Case Study, L1) — Kubernetes Core
- Case Study: DNS Looks Broken — TLS Expired, Fix Is Cert-Manager (Case Study, L2) — Kubernetes Core
- Case Study: DaemonSet Blocks Eviction (Case Study, L2) — Kubernetes Core
- Case Study: Deployment Stuck — ImagePull Auth Failure, Vault Secret Rotation (Case Study, L2) — Kubernetes Core
- Case Study: Drain Blocked by PDB (Case Study, L2) — Kubernetes Core
- Case Study: HPA Flapping — Metrics Server Clock Skew, Fix Is NTP (Case Study, L2) — Kubernetes Core
- Case Study: ImagePullBackOff Registry Auth (Case Study, L1) — Kubernetes Core