---
tags:
- security
- l1
- flashcard-deck
- incident-triage
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Incident Triage](../../../../library/portal/topics.md) | **Domain:** Security
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
incident-triage/a1d2e3f4b5c6	incident-triage	easy	incident-triage, severity, basics	What are the four standard incident severity levels and their meanings?	SEV-1: complete outage or data loss (immediate response). SEV-2: major degradation or partial outage (< 30 min response). SEV-3: minor degradation with workaround (< 4 hours). SEV-4: cosmetic or informational (next business day).	training/library/topics/incident-triage/primer.md
incident-triage/b2e3f4a5c6d7	incident-triage	easy	incident-triage, first-response	What is the first thing you should do when an alert fires?	Acknowledge the alert to stop the escalation timer, then read the alert message and any linked runbook. Check if it is a known issue or repeat incident before diving into diagnosis.	training/library/topics/incident-triage/primer.md
incident-triage/c3f4a5b6d7e8	incident-triage	easy	incident-triage, communication	What information should an incident channel post contain?	Severity level, brief description, current status (Investigating/Identified/Monitoring/Resolved), impact summary, Incident Commander name, last update timestamp, and next update time.\n\nRemember: Mnemonic SSCINL — Severity, Status, Commander, Impact, Next update, Link. Six elements for every incident post.\n\nGotcha: Missing the next-update-time is the #1 communication failure. Silence breeds panic.	training/library/topics/incident-triage/primer.md
incident-triage/d4a5b6c7e8f9	incident-triage	medium	incident-triage, blast-radius	What questions should you ask when assessing blast radius?	Which services are affected? Which regions/zones? How many users impacted? Is it getting worse, stable, or recovering? Are dependent systems at risk? Is there data integrity risk?\n\nRemember: Blast radius assessment = scope the damage before fixing. Mnemonic SWIG-D: Services, Width (regions), Impact (users), Getting worse?, Data integrity.\n\nGotcha: Skip blast radius assessment and you may fix one symptom while the real problem spreads.	training/library/topics/incident-triage/primer.md
incident-triage/e5b6c7d8f9a0	incident-triage	medium	incident-triage, escalation	When should you escalate an incident instead of continuing to troubleshoot alone?	Escalate when: you have spent 15 minutes without understanding the failure mode, the issue is in a system you do not own, severity needs to be raised, additional expertise is required, or customer impact is growing.	training/library/topics/incident-triage/primer.md
incident-triage/f6c7d8e9a0b1	incident-triage	medium	incident-triage, escalation, context	What five pieces of context should you provide when escalating an incident?	1) What is happening (symptoms, not theories). 2) What has been tried so far. 3) When it started. 4) Current blast radius. 5) Links to dashboards, logs, and alerts.\n\nRemember: Escalation context mnemonic WWTBL: What happened, What tried, Time started, Blast radius, Links to dashboards.\n\nGotcha: Never escalate with just it is broken. Context saves the next responder 15+ minutes of re-discovery.	training/library/topics/incident-triage/primer.md
incident-triage/a7d8e9f0b1c2	incident-triage	medium	incident-triage, verification	Why should you verify the alert signal before mobilizing a full response?	Monitoring can produce false positives from misconfigured thresholds, flapping metrics, or stale checks. Confirming from multiple data sources (metrics, logs, synthetic checks) prevents wasting time and team energy on phantom incidents.	training/library/topics/incident-triage/primer.md
incident-triage/b8e9f0a1c2d3	incident-triage	hard	incident-triage, anti-patterns	What is "premature root cause" and why is it dangerous during triage?	Declaring root cause before verifying it leads to fixing symptoms while the real problem continues. It creates tunnel vision where contradicting evidence is ignored, potentially extending the incident.\n\nAnalogy: Like a doctor diagnosing chest pain as heartburn before running an EKG — premature diagnosis can be fatal.\n\nRemember: Correlation is not causation — the deploy happened before the outage, but that does not mean it caused it.	training/library/topics/incident-triage/primer.md
incident-triage/c9f0a1b2d3e4	incident-triage	hard	incident-triage, communication	What are the key communication anti-patterns during an incident?	Blaming individuals during the incident, providing unsupportable ETAs, going silent for more than 30 minutes on SEV-1, sharing technical details with non-technical stakeholders, and forgetting to update the status page.	training/library/topics/incident-triage/primer.md
incident-triage/d0a1b2c3e4f5	incident-triage	hard	incident-triage, post-triage	What should happen after an incident is mitigated but before it is closed?	Write a timeline while memory is fresh, keep the incident channel open for follow-up, schedule a blameless postmortem within 48 hours for SEV-1/2, track action items to completion, and update runbooks if they were inadequate.	training/library/topics/incident-triage/primer.md
incident-triage/e1b2c3d4f5a6	incident-triage	medium	incident-triage, roles, handoff	When and how should the Incident Commander role be transferred during a long-running incident?	Transfer IC when the current IC is fatigued (2+ hours), when a shift boundary is reached, or when a subject-matter expert should lead. The outgoing IC briefs the incoming IC on status, timeline, blast radius, and next actions, then announces the handoff in the incident channel.	
incident-triage/f2c3d4e5a6b7	incident-triage	hard	incident-triage, rollback, decision	How do you decide between rolling back a change and pushing a forward-fix during an incident?	Roll back when: the failing change is identified with high confidence, rollback is tested and fast (< 5 min), and data integrity is not at risk. Forward-fix when: rollback is impossible (schema migration, data change), the fix is small and well-understood, or rollback would cause equal or worse impact.	
incident-triage/a3d4e5f6b7c8	incident-triage	hard	incident-triage, communication, customer	What should customer-facing communication include during a SEV-1 incident?	Acknowledge the issue without speculating on cause, state the known impact, provide an estimated next-update time (not ETA to resolution), use plain language without internal jargon, and update at regular intervals (every 30 min for SEV-1) even if status has not changed.	
incident-triage/b4e5f6a7c8d9	incident-triage	medium	incident-triage, timeline, documentation	What should an incident timeline capture, and when should you start writing it?	Start the timeline immediately when the incident is declared. Record: alert fire time, first responder actions, each escalation, key diagnostic findings, mitigation attempts (successful and failed), and resolution confirmation. Timestamps should be in UTC.	
incident-triage/c5f6a7b8d9e0	incident-triage	medium	incident-triage, containment, blast-radius	What are practical blast radius containment techniques during an active incident?	Disable the feature flag that triggered the issue, shift traffic away from the affected region (DNS or load balancer), scale down the offending service, block the problematic endpoint at the gateway, or isolate the affected database replica. Goal: stop the bleeding before diagnosing root cause.	
incident-triage/d6a7b8c9e0f1	incident-triage	easy	incident-triage, preparation, runbook	What makes a runbook effective during an incident versus one that gets ignored?	Effective runbooks have: a clear trigger condition (when to use this), step-by-step commands (copy-pasteable), expected output for each step, escalation contacts with current names, and a last-verified date. Runbooks older than 90 days without verification are unreliable.	
incident-triage/f39ed8b67a20	incident-triage	medium	triage;severity	What are the typical severity levels in an incident management system?	SEV1 (critical, customer-facing outage), SEV2 (major degradation, partial outage), SEV3 (minor issue, limited impact), SEV4/SEV5 (informational, cosmetic). Each level triggers different response expectations.\n\nRemember: SEV1=all-hands, SEV2=team-response, SEV3=next-sprint, SEV4/5=backlog. Response time doubles with each level.\n\nGotcha: Under-declaring severity delays response. Over-declaring causes alert fatigue. When in doubt, declare higher and downgrade.	
incident-triage/d9f146ba0b03	incident-triage	hard	triage;blast-radius	How do you assess blast radius during the first 5 minutes of an incident?	Check: which services are affected (dependency graph), which customers are impacted (traffic/error dashboards), is the issue spreading (error rate trend), and what changed recently (deploy log, config changes).\n\nRemember: First 5 minutes: dependency graph + error dashboards + deploy log + error trend. Four data sources in parallel.\n\nGotcha: Is it spreading? is the most critical question. A growing blast radius means containment is priority #1.	
incident-triage/7f6e629a919e	incident-triage	medium	triage;communication	Why should you declare an incident early rather than waiting to confirm the issue?	Declaring early enables coordination and visibility before the problem grows. False alarms are cheap; delayed response to real incidents is expensive. 'Declare first, investigate second' reduces MTTR.\n\nRemember: Declare first, investigate second. The cost of a false alarm is low; the cost of a delayed response is high.\n\nAnalogy: Like pulling a fire alarm — better to evacuate for a false alarm than to delay during a real fire.	
incident-triage/69268638a3bf	incident-triage	medium	triage;runbooks	What makes a runbook useful during incident triage?	A good runbook provides: symptom-to-action mapping, diagnostic commands to run, escalation paths, and known fixes. It reduces cognitive load and ensures consistent response regardless of who is on-call.\n\nRemember: Good runbook = symptom-to-action mapping. Bad runbook = 50-page document nobody reads during a 3 AM outage.\n\nGotcha: Untested runbooks are worse than no runbook — they give false confidence. Verify quarterly.	

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Incident Triage](../../../../library/topics/incident-triage/index.md) (Topic Pack, L1) — Incident Triage
- [Runbook: CVE Response (Critical Vulnerability)](../../../../library/runbooks/security/cve-response.md) (Runbook, L2) — Incident Triage
- [Runbook: Unauthorized Access Investigation](../../../../library/runbooks/security/unauthorized-access.md) (Runbook, L2) — Incident Triage

<!-- wiki:related:end -->
