Skip to content

Portal | Level: L2: Operations | Topics: Incident Response | Domain: Kubernetes

Guided Investigation Engine

A CLI-based coaching tool that teaches systematic SRE debugging. It pairs with the existing incident/chaos tools to provide structured investigation playbooks, progressive hints, and a journal for recording findings.

Quick Start

# 1. Deploy the stack
make deploy-all

# 2. Inject an incident (via chaos scripts or manual trigger)
training/interactive/investigation/trigger.sh probe-failure
# or: training/interactive/chaos/scripts/break_readiness.sh --yes

# 3. Investigate
make investigate        # See step-by-step plan
make hint               # Get hint level 1
make hint HINT=2        # Deeper hint

# 4. Fix the issue using kubectl/helm

# 5. Record what you found
make explain            # Generate journal template

# 6. Clear the incident
training/interactive/investigation/trigger.sh --clear

Commands

Command What it does
make investigate Print investigation plan for the active incident
make hint Show progressive hint (HINT=1..4)
make explain Create/append a journal entry
make investigate-selftest Run self-tests
training/interactive/investigation/trigger.sh <id> Manually trigger an incident
training/interactive/investigation/trigger.sh --list List known incident IDs
training/interactive/investigation/trigger.sh --clear Clear active incident

Playbooks

12 playbooks covering common Kubernetes failure modes:

Playbook Covers
service-down Generic service unavailability (fallback)
crashloopbackoff Container crash loops
probe-failure Readiness/liveness probe failures
imagepullbackoff Image pull errors
oomkilled Out-of-memory kills
hpa-not-scaling HPA not responding to load
dns-resolution DNS resolution failures
networkpolicy-block NetworkPolicy blocking traffic
helm-upgrade-failed Failed Helm upgrades
prometheus-target-down Prometheus scrape target down
loki-no-logs Missing logs in Loki
latency-spike High latency investigation

Each playbook links to existing runbooks, labs, and Quest Ladder exercises. See map.tsv for the full mapping.

Hint System

Hints are progressive (1-4) and avoid spoilers: 1. Where to look (which commands to run) 2. Narrows the subsystem (probe/service/selector/etc.) 3. Points at the likely misconfiguration 4. Near-spoiler with fix direction

Journal

Entries are written to training/interactive/investigation/journal/<date>-<incident_id>.md. Use make explain to generate a template, then fill in: - Symptom - Key evidence (3 bullets) - Root cause - Fix commands - Verification commands

File Structure

training/interactive/investigation/
├── README.md           # This file
├── map.tsv             # incident -> playbook/runbook/lab/quest mapping
├── investigate.sh      # Main investigation script
├── hints.sh            # Progressive hint script
├── explain.sh          # Journal entry generator
├── trigger.sh          # Manual incident trigger
├── selftest.sh         # Self-test suite
├── lib/
│   └── common.sh       # Shared functions
├── playbooks/          # 12 investigation playbooks
├── hints/              # Progressive hint files (per incident)
└── journal/            # Investigation journal entries

Wiki Navigation

Prerequisites

  • Incident Simulator (18 scenarios) (CLI) (Exercise Set, L2)