Generic Hints¶
Fallback hints when no incident-specific hint file exists.
Hint 1¶
Start broad: check pod status, events, and logs.
- kubectl get pods -n grokdevops
- kubectl get events -n grokdevops --sort-by='.lastTimestamp' | tail -15
- kubectl logs -n grokdevops <pod> --tail=20
Hint 2¶
Narrow down: is the issue at the pod level, service level, or infrastructure level?
- Pod: check kubectl describe pod for container state, exit codes, probe status
- Service: check kubectl get endpoints for missing backends
- Infra: check kubectl top nodes and kubectl describe node for pressure
Hint 3¶
Common root causes to check: - Wrong image tag or missing image -> check deployment image field - Broken probe path -> check readinessProbe/livenessProbe config - Resource limits too low -> check resources.limits in deployment - Missing ConfigMap/Secret -> check volume mounts and envFrom - Label mismatch -> compare Service selector with Pod labels
Hint 4¶
If still stuck:
- helm get values grokdevops -n grokdevops to see current config
- helm diff (if installed) or compare with devops/helm/values-dev.yaml
- Read the matching runbook in training/library/runbooks/
- Check training/library/guides/troubleshooting.md for known issues