Portal | Level: L2: Operations | Topics: Incident Response | Domain: Kubernetes
Guided Investigation Engine¶
A CLI-based coaching tool that teaches systematic SRE debugging. It pairs with the existing incident/chaos tools to provide structured investigation playbooks, progressive hints, and a journal for recording findings.
Quick Start¶
# 1. Deploy the stack
make deploy-all
# 2. Inject an incident (via chaos scripts or manual trigger)
training/interactive/investigation/trigger.sh probe-failure
# or: training/interactive/chaos/scripts/break_readiness.sh --yes
# 3. Investigate
make investigate # See step-by-step plan
make hint # Get hint level 1
make hint HINT=2 # Deeper hint
# 4. Fix the issue using kubectl/helm
# 5. Record what you found
make explain # Generate journal template
# 6. Clear the incident
training/interactive/investigation/trigger.sh --clear
Commands¶
| Command | What it does |
|---|---|
make investigate |
Print investigation plan for the active incident |
make hint |
Show progressive hint (HINT=1..4) |
make explain |
Create/append a journal entry |
make investigate-selftest |
Run self-tests |
training/interactive/investigation/trigger.sh <id> |
Manually trigger an incident |
training/interactive/investigation/trigger.sh --list |
List known incident IDs |
training/interactive/investigation/trigger.sh --clear |
Clear active incident |
Playbooks¶
12 playbooks covering common Kubernetes failure modes:
| Playbook | Covers |
|---|---|
| service-down | Generic service unavailability (fallback) |
| crashloopbackoff | Container crash loops |
| probe-failure | Readiness/liveness probe failures |
| imagepullbackoff | Image pull errors |
| oomkilled | Out-of-memory kills |
| hpa-not-scaling | HPA not responding to load |
| dns-resolution | DNS resolution failures |
| networkpolicy-block | NetworkPolicy blocking traffic |
| helm-upgrade-failed | Failed Helm upgrades |
| prometheus-target-down | Prometheus scrape target down |
| loki-no-logs | Missing logs in Loki |
| latency-spike | High latency investigation |
Each playbook links to existing runbooks, labs, and Quest Ladder exercises. See map.tsv for the full mapping.
Hint System¶
Hints are progressive (1-4) and avoid spoilers: 1. Where to look (which commands to run) 2. Narrows the subsystem (probe/service/selector/etc.) 3. Points at the likely misconfiguration 4. Near-spoiler with fix direction
Journal¶
Entries are written to training/interactive/investigation/journal/<date>-<incident_id>.md. Use make explain to generate a template, then fill in:
- Symptom
- Key evidence (3 bullets)
- Root cause
- Fix commands
- Verification commands
File Structure¶
training/interactive/investigation/
├── README.md # This file
├── map.tsv # incident -> playbook/runbook/lab/quest mapping
├── investigate.sh # Main investigation script
├── hints.sh # Progressive hint script
├── explain.sh # Journal entry generator
├── trigger.sh # Manual incident trigger
├── selftest.sh # Self-test suite
├── lib/
│ └── common.sh # Shared functions
├── playbooks/ # 12 investigation playbooks
├── hints/ # Progressive hint files (per incident)
└── journal/ # Investigation journal entries
Wiki Navigation¶
Prerequisites¶
- Incident Simulator (18 scenarios) (CLI) (Exercise Set, L2)
Related Content¶
- Change Management (Topic Pack, L1) — Incident Response
- Chaos Engineering Scripts (CLI) (Exercise Set, L2) — Incident Response
- Debugging Methodology (Topic Pack, L1) — Incident Response
- Incident Command & On-Call (Topic Pack, L2) — Incident Response
- Incident Response Flashcards (CLI) (flashcard_deck, L1) — Incident Response
- Incident Simulator (18 scenarios) (CLI) (Exercise Set, L2) — Incident Response
- Ops War Stories & Pattern Recognition (Topic Pack, L2) — Incident Response
- Postmortems & SLOs (Topic Pack, L2) — Incident Response
- Runbook Craft (Topic Pack, L1) — Incident Response
- Systems Thinking for Engineers (Topic Pack, L1) — Incident Response