---
tags:
- devops
- l1
- flashcard-deck
- incident-response
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Incident Response](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
incident-response/a1b2c3d4e5f6	incident-response	easy	incident-response,roles	What is the Incident Commander (IC) role?	The IC owns the incident lifecycle: declares severity, coordinates responders, manages communication, and decides when to escalate or resolve. The IC does NOT debug — they facilitate. They run the war room, keep a timeline, and ensure status updates go out on schedule.	training/library/topics/incident-response/primer.md
incident-response/b2c3d4e5f6a7	incident-response	easy	incident-response,communication	What should an initial incident notification include?	Severity level, affected service(s), user impact summary, current status (Investigating/Identified/Monitoring/Resolved), who is responding, next update ETA, and a link to the incident channel or bridge call. Over-communicate early — silence breeds confusion.	training/library/topics/incident-response/primer.md
incident-response/c3d4e5f6a7b8	incident-response	medium	incident-response,rollback	When should you roll back vs fix forward?	Roll back when: the cause is a recent deploy, rollback is safe and fast, no irreversible data migrations have run. Fix forward when: rollback would cause data loss, the fix is small and well-understood, or the issue predates the last deploy. Default to rollback if unsure — speed matters in outages.	training/library/topics/incident-response/primer.md
incident-response/d4e5f6a7b8c9	incident-response	medium	incident-response,triage	What are the first three things to check during a production outage?	1) Monitoring dashboards — scope the blast radius (total vs partial, which regions/services). \n2) Recent changes — check deploy logs, config changes, infra changes in the last hour. \n3) Cluster/infra health — node status, pod health, external dependency status. Do NOT start fixing until you understand the scope.	training/library/topics/incident-response/primer.md
incident-response/e5f6a7b8c9d0	incident-response	hard	incident-response,postmortem	What makes a good blameless postmortem?	Focus on systemic causes, not individual mistakes. Include: timeline of events, what went well, what went poorly, root cause analysis (5 Whys), action items with owners and deadlines. Share widely. The goal is learning and prevention, not punishment. Track action item completion.	training/library/topics/incident-response/primer.md
incident-response/f6a7b8c9d0e1	incident-response	medium	incident-response,severity	How do you determine incident severity?	SEV-1: complete outage or data loss/breach. SEV-2: major degradation, significant user impact. SEV-3: minor degradation with workaround available. SEV-4: cosmetic or low-impact. Key factors: number of users affected, revenue impact, data integrity risk, and whether a workaround exists.	training/library/topics/incident-response/primer.md
incident-response/a7b8c9d0e1f2	incident-response	hard	incident-response,coordination	How do you manage multiple simultaneous incidents?	Assign separate Incident Commanders for each. Determine if incidents are related (common root cause). If related, merge into one incident with increased severity. If independent, ensure responders are not overloaded across both. Prioritize by severity — SEV-1 gets resources first.	training/library/topics/incident-response/primer.md
incident-response/b8c9d0e1f2a3	incident-response	easy	incident-response,runbook	What is a runbook and what should it contain?	A runbook is a step-by-step guide for handling a specific alert or operational task. It should include: alert description, impact assessment, diagnostic commands, remediation steps, escalation path, and verification steps. Keep runbooks in version control, link them from alerts, and update after every incident where the runbook was insufficient.	training/library/topics/incident-response/primer.md
incident-response/c9d0e1f2a3b4	incident-response	medium	incident-response,escalation	When should you escalate an incident?	Escalate when: you've been troubleshooting for 15+ minutes without progress, the issue is outside your domain expertise, severity is increasing, customer/business impact is growing, or you need access/permissions you don't have. Escalating early is not failure — it is responsible incident management.	training/library/topics/incident-response/primer.md
incident-response/d0e1f2a3b4c5	incident-response	hard	incident-response,chaos	How does chaos engineering improve incident response?	Chaos engineering (GameDays, Chaos Monkey) intentionally injects failures in controlled conditions to test detection, alerting, runbooks, and team response. It reveals gaps before real incidents do: missing alerts, unclear runbooks, slow escalation paths, and single points of failure. Run regularly and track improvements.	training/library/topics/incident-response/primer.md
incident-response/e1f2a3b4c5d6	incident-response	medium	incident-response,timeline	Why is maintaining an incident timeline critical?	The timeline captures what happened, when, and what actions were taken. During the incident, it prevents duplicate work and keeps new responders oriented. After the incident, it is the foundation for the postmortem and helps identify delays in detection, response, and resolution. Use a shared doc or incident tool — never rely on memory.	training/library/topics/incident-response/primer.md
incident-response/f2a3b4c5d6e7	incident-response	easy	incident-response,statuspage	What is the purpose of a status page during an incident?	A status page (Statuspage.io, Cachet) communicates outage status to customers and internal stakeholders, reducing support ticket volume and building trust through transparency. Update it at every severity change and at regular intervals. Include: affected components, current status, ETA if known, and workarounds.	training/library/topics/incident-response/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Change Management](../../../../library/topics/change-management/index.md) (Topic Pack, L1) — Incident Response
- Chaos Engineering Scripts *(CLI)* (Exercise Set, L2) — Incident Response
- [Debugging Methodology](../../../../library/topics/debugging-methodology/index.md) (Topic Pack, L1) — Incident Response
- [Incident Command & On-Call](../../../../library/topics/incident-command/index.md) (Topic Pack, L2) — Incident Response
- Incident Simulator (18 scenarios) *(CLI)* (Exercise Set, L2) — Incident Response
- Investigation Engine *(CLI)* (Exercise Set, L2) — Incident Response
- [Ops War Stories & Pattern Recognition](../../../../library/topics/ops-war-stories/index.md) (Topic Pack, L2) — Incident Response
- [Postmortems & SLOs](../../../../library/topics/postmortem-slo/index.md) (Topic Pack, L2) — Incident Response
- [Runbook Craft](../../../../library/topics/runbook-craft/index.md) (Topic Pack, L1) — Incident Response
- [Systems Thinking for Engineers](../../../../library/topics/systems-thinking/index.md) (Topic Pack, L1) — Incident Response

<!-- wiki:related:end -->
