---
tags:
- devops
- l1
- flashcard-deck
- on-call
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [On-Call & Incident Command](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
on-call/a1b2c3d4e5f3	on-call	easy	on-call, incident-commander	What is the Incident Commander (IC) role, and what should they NOT do during an incident?	The IC coordinates the response: declares the incident and severity, opens the war room, assigns roles, sets investigation direction, decides on escalation, and declares resolution. The IC does NOT debug code, SSH into servers, write queries, or get tunnel-visioned on one theory.\n\nRemember: Detect→Triage→Mitigate→Resolve→Postmortem. "DTMRP."	training/library/topics/incident-command/primer.md
on-call/b2c3d4e5f3a4	on-call	easy	on-call, roles	What are the four key incident response roles and their responsibilities?	Incident Commander (coordinates response, makes decisions), Technical Lead (drives investigation and remediation), Communications Lead (updates statuspage, customers, stakeholders), and Scribe (records timeline, decisions, and actions in real-time).\n\nRemember: Detect→Triage→Mitigate→Resolve→Postmortem. "DTMRP."	training/library/topics/incident-command/primer.md
on-call/c3d4e5f3a4b5	on-call	easy	on-call, severity	How do you decide between SEV-1 and SEV-2 using a severity decision tree?	If the service is completely down, it is SEV-1. If not completely down but more than 10% of users are affected without a workaround, it is SEV-2. If fewer users are affected, it is SEV-3. No customer impact is SEV-4.\n\nRemember: SEV1=outage, SEV2=degraded, SEV3=limited, SEV4=cosmetic.	training/library/topics/incident-command/primer.md
on-call/d4e5f3a4b5c6	on-call	medium	on-call, pagerduty, escalation	What is a typical PagerDuty escalation policy structure and why are escalation timeouts important?	Level 1: Primary on-call (5-min ack window). Level 2: Secondary on-call (5-min ack). Level 3: Engineering manager (10-min ack). Level 4: VP Engineering (phone call). Escalation timeouts prevent a missed acknowledgment from blocking the entire incident response.\n\nRemember: primary→secondary→manager→VP. Tools: PagerDuty, Opsgenie, VictorOps.	training/library/topics/incident-command/primer.md
on-call/e5f3a4b5c6d7	on-call	medium	on-call, handoff	What must be included in an on-call handoff, and why is a written handoff non-negotiable?	A handoff must include: active issues (with ticket references and runbooks), recent infrastructure changes, watch items (upcoming events), and known noisy alerts. Written handoffs are non-negotiable because the outgoing engineer disappearing without context leaves the incoming engineer blind to ongoing issues.\n\nRemember: Good on-call = runbooks + escalation + blameless postmortems.\n\nGotcha: Every alert should be actionable. Can wait until morning? Not page-worthy.	training/library/topics/incident-command/primer.md
on-call/f3a4b5c6d7e8	on-call	medium	on-call, runbooks	What should a runbook contain and why is it important for on-call?	A runbook should contain: symptoms (what alerts fire, what users see), diagnosis steps (specific dashboards and commands), remediation steps (detailed with exact commands), and escalation paths (when and who to escalate to). Runbooks let any on-call engineer handle an issue at 3am without prior knowledge of the service.\n\nRemember: Runbooks = step-by-step guides. Include: check, mitigate, escalate. Keep updated.	training/library/topics/incident-command/primer.md
on-call/04a5b6c7d8e9	on-call	medium	on-call, communication	Why should incident updates be sent on a regular cadence even when there is no new information?	Stakeholders who hear nothing for an hour assume the worst and may interfere with the response. Setting a timer for updates every 15-30 minutes, even if the update is "still investigating," maintains trust and keeps stakeholders informed without them needing to ask.\n\nRemember: Detect→Triage→Mitigate→Resolve→Postmortem. "DTMRP."	training/library/topics/incident-command/primer.md
on-call/15b6c7d8e9f0	on-call	hard	on-call, burnout, health	What on-call health metrics should you track monthly, and what are the red flags?	Track: pages per shift (target <2 per day, 0 per night), time-to-ack (<5 min), time-to-resolve (<30 min for P1), sleep interruptions, and satisfaction score (target >3.5/5). Red flags: pages trending up, same alerts recurring weekly, satisfaction below 3.0, or an engineer requesting permanent removal from rotation.\n\nRemember: Good on-call = runbooks + escalation + blameless postmortems.\n\nGotcha: Every alert should be actionable. Can wait until morning? Not page-worthy.	training/library/topics/incident-command/primer.md
on-call/26c7d8e9f0a1	on-call	hard	on-call, rotation, design	What is the recommended on-call rotation design, and why should the secondary always be next week's primary?	Use 1-week shifts (Mon 09:00 to Mon 09:00) with a 30-minute handoff overlap and written summary. The secondary is always next week's primary so they gain context before becoming primary. This ensures the incoming primary has already seen the previous week's issues as secondary backup.\n\nRemember: Good on-call = runbooks + escalation + blameless postmortems.\n\nGotcha: Every alert should be actionable. Can wait until morning? Not page-worthy.	training/library/topics/incident-command/primer.md
on-call/37d8e9f0a1b2	on-call	hard	on-call, escalation, pitfalls	Why should escalation be reframed as "the system working correctly" rather than as a failure?	Engineers often wait too long to escalate because they fear looking incompetent. This delay extends incidents and increases blast radius. Reframing escalation as the system working correctly encourages timely escalation. If escalation to Level 4 (VP Engineering) happens regularly, the problem is the L1-L3 process, not the people.\n\nRemember: primary→secondary→manager→VP. Tools: PagerDuty, Opsgenie, VictorOps.	training/library/topics/incident-command/primer.md
on-call/48e9f0a1b2c3	on-call	easy	on-call, rotation, scheduling	What are the key principles for designing a fair on-call rotation?	1-week shifts (not 2 — burnout risk). Equal distribution across team members. Minimum 2 people per rotation (primary + secondary). Allow shift swaps with advance notice. Avoid scheduling during planned PTO. Ensure timezone coverage for global teams. Secondary becomes next week's primary for context continuity. Track on-call load per person quarterly and rebalance if uneven.\n\nRemember: Good on-call = runbooks + escalation + blameless postmortems.\n\nGotcha: Every alert should be actionable. Can wait until morning? Not page-worthy.	
on-call/59f0a1b2c3d4	on-call	medium	on-call, escalation, policies	What should an escalation policy define and what are common mistakes?	Define: who is paged at each level, ack timeout before escalation (typically 5 min), maximum escalation depth (usually 4 levels), and after-hours behavior. Common mistakes: no secondary on-call (single point of failure), ack timeouts too long (15+ minutes delays response), escalating to managers who cannot fix technical issues, and not testing the escalation path regularly. Test by running a drill page monthly.\n\nRemember: primary→secondary→manager→VP. Tools: PagerDuty, Opsgenie, VictorOps.	
on-call/60a1b2c3d4e5	on-call	medium	on-call, handoff, procedures	What are the steps for an effective on-call handoff?	1) Outgoing engineer writes a handoff document: active incidents, recent deployments, watch items, known noisy alerts.\n2) Overlap meeting (15-30 min) to walk through the document.\n3) Verify incoming engineer has access to all tools (PagerDuty, dashboards, VPN).\n4) Transfer the PagerDuty rotation at the agreed time.\n5) Outgoing remains available for 1 hour post-handoff for questions.\nNever rely on verbal-only handoffs — written documentation is required.\n\nRemember: Good on-call = runbooks + escalation + blameless postmortems.\n\nGotcha: Every alert should be actionable. Can wait until morning? Not page-worthy.	
on-call/71b2c3d4e5f6	on-call	easy	on-call, fatigue, prevention	What are the warning signs of on-call fatigue and how do you address them?	Warning signs: ignoring non-critical alerts, delayed ack times trending up, engineers requesting removal from rotation, increased sick days during on-call weeks, and cynicism about alerting. Address by: reducing alert noise (fix or delete noisy alerts), ensuring no more than 2 pages per day-shift, zero pages as the night-shift goal, compensating on-call time, and rotating fairly so no one is on-call more than 25% of the time.\n\nRemember: Alert fatigue → missed alerts. Cure: delete noise, aggregate, proper thresholds.	
on-call/82c3d4e5f6a7	on-call	medium	on-call, severity, classification	How should you classify incident severity and what response does each level require?	SEV-1 (Critical): total service outage or data loss, all-hands response, 5-min ack, external comms within 15 min.\nSEV-2 (Major): significant degradation affecting >10% of users, dedicated IC + tech lead, 15-min ack.\nSEV-3 (Minor): limited impact with workaround available, primary on-call handles, 30-min ack.\nSEV-4 (Low): cosmetic or no customer impact, track as a ticket. When in doubt, declare higher severity — it is easier to downgrade than to escalate late.\n\nRemember: Detect→Triage→Mitigate→Resolve→Postmortem. "DTMRP."	
on-call/93d4e5f6a7b8	on-call	hard	on-call, post-incident, review	When should a post-incident review be triggered and what must it cover?	Trigger for: all SEV-1 and SEV-2 incidents, any incident lasting >1 hour, incidents requiring escalation beyond L2, and near-misses that could have been severe. Must cover: timeline of events, root cause analysis (not blame), what detection/response worked, what failed, and concrete action items with owners and due dates. Hold the review within 48 hours while memory is fresh. Action items without owners and deadlines are worthless.\n\nRemember: Detect→Triage→Mitigate→Resolve→Postmortem. "DTMRP."	
on-call/04e5f6a7b8c9	on-call	easy	on-call, compensation, fairness	What are common approaches to on-call compensation?	Flat stipend per on-call shift (e.g., $200-500/week). Per-page bonus on top of the stipend. Time-off-in-lieu (comp day after a disruptive shift). Reduced sprint velocity during on-call weeks. Some companies combine stipend + per-page + comp time for severe incidents. The key principle: on-call is real work outside normal hours and must be compensated. Uncompensated on-call leads to burnout and attrition.\n\nRemember: Good on-call = runbooks + escalation + blameless postmortems.\n\nGotcha: Every alert should be actionable. Can wait until morning? Not page-worthy.	
on-call/15f6a7b8c9d0	on-call	hard	on-call, runbooks, maintenance	How do you keep runbooks accurate and useful over time?	Link every alert to its runbook (runbook_url annotation in Prometheus). Review and update runbooks during post-incident reviews. Assign runbook ownership to the service-owning team. Include a "last verified" date at the top — stale runbooks are worse than no runbook (they waste time with wrong commands). Test runbooks during game days. Archive runbooks for decommissioned services. Use version control (git) so changes are tracked and reviewable.\n\nRemember: Runbooks = step-by-step guides. Include: check, mitigate, escalate. Keep updated.	
on-call/90a9f4a6f2e4	on-call	medium	on-call;escalation	What is the purpose of an escalation policy in an on-call rotation?	It defines when and how to escalate an unacknowledged or unresolved alert to the next responder or team, ensuring incidents don't stall.\n\nRemember: primary→secondary→manager→VP. Tools: PagerDuty, Opsgenie, VictorOps.	
on-call/2896adaf5178	on-call	hard	on-call;fatigue	How does alert fatigue reduce incident response quality, and what is one structural countermeasure?	Alert fatigue causes responders to ignore or slow-respond to alerts. Countermeasure: tune alert thresholds and suppress duplicate/flapping alerts so only actionable alerts fire.\n\nRemember: Detect→Triage→Mitigate→Resolve→Postmortem. "DTMRP."	

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Incident Command & On-Call](../../../../library/topics/incident-command/index.md) (Topic Pack, L2) — On-Call & Incident Command
- [On-Call](../../../../library/topics/on-call/index.md) (Topic Pack, L2) — On-Call & Incident Command
- [Runbook Craft](../../../../library/topics/runbook-craft/index.md) (Topic Pack, L1) — On-Call & Incident Command
- [The Psychology of Incidents](../../../../library/topics/incident-psychology/index.md) (Topic Pack, L2) — On-Call & Incident Command
- [Vendor Management & Escalation](../../../../library/topics/vendor-management/index.md) (Topic Pack, L1) — On-Call & Incident Command

<!-- wiki:related:end -->
