top of page

Incident Response Runbook Template

Shawn West
Jun 21
5 min read

Updated: 2 days ago

A runbook is a written guide for handling a specific incident type. It exists because at 3am, the on-call engineer needs to know what to do, not what to think about. Good runbooks are short, scannable, and tested. Bad runbooks are long, narrative, and stale.


This guide is a template optimized for the 3am reality.


The Working Template


# Runbook: [Specific Incident Type]
Last verified: YYYY-MM-DD
Owner: [Team/role]
Severity: [P0/P1/P2 — when this applies]

## Symptom
[What you see that triggers this runbook]

## Quick Triage (first 5 minutes)
1. [Specific check]
2. [Specific check]
3. [Specific check]

## Diagnosis Tree
- If [symptom A] → see [Mitigation 1]
- If [symptom B] → see [Mitigation 2]
- If [symptom C] → see [Escalation]

## Mitigation 1: [Common cause name]
1. [Exact command or action]
2. [Exact command or action]
3. [Verification step]

## Mitigation 2: [Another cause]
[Same structure]

## Escalation
- Page: [specific person/role]
- Channel: [specific channel]
- Context to share: [what to communicate]

## Verification
[How to confirm the issue is resolved]

## After-action
[What to capture for the post-incident review]

The structure is action-oriented. Each section answers a specific question at a specific moment.


Symptom Description


The first thing the on-call sees should help them confirm they're in the right runbook.


Bad: "Database issues."


Good: "Checkout 5xx errors > 1% for 5+ minutes; database connection errors in logs."


Specific symptoms include observable signals — error rates, log patterns, alert names. The on-call confirms "yes, this matches what I'm seeing" and proceeds with confidence.


Quick Triage


The first five minutes of an incident are critical. The runbook should immediately give the on-call useful actions.


Examples:


  • "Check the dashboard at [link]; note the affected services"

  • "Run kubectl get pods -n prod and identify any pods in CrashLoopBackOff"

  • "Verify that the alert was real by checking [specific metric]"


These are fast checks that narrow the problem. They're not the fix — they're the orientation.


Diagnosis Tree


After triage, the runbook branches based on what the on-call sees.


- If "connection refused" in logs → see Mitigation 1 (service down)
- If "timeout" in logs → see Mitigation 2 (overload)
- If logs are clean → see Mitigation 3 (external dependency)
- If none of these match → see Escalation

The tree lets the on-call navigate quickly. Each branch leads to a specific mitigation.


Mitigations


Specific actions to resolve the most common causes.


Format:


  • Numbered steps

  • Exact commands (not descriptions of commands)

  • Verification after each step

  • Time estimates if relevant


Example:


## Mitigation 1: Service is down (connection refused)

1. Check pod status: `kubectl get pods -n prod -l app=checkout`
2. If pods are crashing, view recent logs: `kubectl logs -n prod -l app=checkout --tail=200`
3. Restart the deployment: `kubectl rollout restart deployment/checkout -n prod`
4. Wait for pods to become Ready (typically 30-60s)
5. Verify health endpoint: `curl -f https://checkout.internal/health`
6. Check error rate dashboard returns to baseline (typically 2-3 minutes)

The on-call can execute this without thinking. The runbook does the thinking.


Escalation


When the runbook can't resolve the issue.


  • Who to page: specific name or role

  • Where to escalate: specific channel or thread

  • What to include: the context the next person needs


Don't say "escalate to the appropriate team." Say "page the database on-call (role: db-oncall in PagerDuty)."


Verification


How to confirm the issue is resolved. Specific signals to check.


  • Dashboard metrics returned to baseline (link)

  • Alert cleared in alerting system

  • Test transaction completes successfully

  • No new error reports in support channel


The on-call should not declare an incident resolved on intuition. Verification is the evidence.


After-Action


What to do once the incident is over.


  • Update the incident channel with the final summary

  • Notify any affected stakeholders

  • Capture timeline for the post-mortem

  • File ticket for any underlying fix

  • Update the runbook if anything was missing


The after-action section closes the loop. Without it, the incident ends but the learning doesn't get captured.


What Makes Runbooks Useful


Specific commands, not descriptions. "Run kubectl restart" is worse than the exact command. Copy-pasteable beats description.


Tested. A runbook that's never been executed is fiction. Test it during game days; correct it after real incidents.


Recent. "Last verified" date at the top. If it's more than 6 months old, the on-call should treat it skeptically.


Owned. A specific team or role keeps it current. Without an owner, it goes stale.


Linked. From dashboards, alerts, and the incident-response toolchain. A runbook no one can find isn't useful.


What Makes Runbooks Useless


Narrative. Paragraphs about how the system works. Useful in onboarding docs; useless at 3am.


Generic. Steps so general they could apply to anything. Specific or skip.


Stale. References to systems, tools, or people that don't exist anymore.


Unverified. Steps that no one has actually run. Often wrong in subtle ways.


Single-author. One person knows it works; everyone else gets stuck.


A Worked Example


# Runbook: Checkout 5xx Error Spike
Last verified: 2026-04-12
Owner: payments-oncall
Severity: P0 when > 5%, P1 when 1-5%

## Symptom
Checkout 5xx error rate > 1% for 5+ minutes. Alert: ALERT-CHECKOUT-5XX.

## Quick Triage (first 5 minutes)
1. Open the checkout dashboard: https://grafana.internal/d/checkout
2. Note: error rate, request volume, latency
3. Check #incidents channel for any active incidents
4. Check status page for upstream provider issues (Stripe, AWS)

## Diagnosis Tree
- If errors are 500s with "database" in logs → Mitigation 1
- If errors are 502/503 → Mitigation 2 (load balancer / pod issues)
- If errors are 504 (timeout) → Mitigation 3 (overload or slow dependency)
- If status page shows AWS issue → see AWS Regional Failover runbook
- If none match → Escalation

## Mitigation 1: Database connection errors
1. Check DB connection pool: `curl checkout.internal/admin/db-stats`
2. If pool exhausted, restart deployment: `kubectl rollout restart deployment/checkout -n prod`
3. Monitor recovery on dashboard
4. If issue persists after restart → check DB itself: see DB runbook

## Mitigation 2: 502/503 from load balancer
1. Check pod status: `kubectl get pods -n prod -l app=checkout`
2. If pods crashing, view logs: `kubectl logs -n prod -l app=checkout --tail=200`
3. Restart deployment if needed
4. Verify pods come Ready
5. Confirm load balancer health checks passing

## Mitigation 3: 504 timeouts
1. Check downstream services on dashboard: payment-service, inventory-service
2. If slow, page their respective on-calls
3. Enable circuit breaker: `kubectl annotate deployment/checkout cb=on`
4. Monitor recovery

## Escalation
- Page: senior-checkout-engineer (role: checkout-senior-oncall)
- Channel: #incident-active
- Context: symptom, what you've tried, current state

## Verification
- Error rate < 0.1% on dashboard
- Successful test transaction: `./scripts/test-checkout.sh`
- No new reports in #support-internal

## After-action
- Update #incidents with summary and resolution
- File POSTMORTEM ticket if customer-impacting > 5 minutes
- Update runbook if any step was unclear or missing

A 3am on-call can navigate this. The information density is high and the actions are concrete.


Maintaining Runbooks


Runbooks rot. Maintenance practices:


  • After every real incident: update the runbook if anything was missing or wrong. The freshest perspective.

  • Quarterly review: owner walks the runbooks, checks for staleness.

  • Game day testing: periodically execute runbooks against staging or simulated incidents. Finds the broken parts.

  • Onboarding test: new on-call walks the runbook before going on rotation. Catches assumed knowledge.


Per-Incident vs. General Runbooks


Two types:


Per-incident: "Checkout 5xx error spike" — specific symptom, specific actions.


General: "How to handle any P0 incident" — process for incidents broadly.


Both are useful. The per-incident runbooks are most-used and most-valuable. The general runbook is the meta-framework.


Key Takeaway


A runbook is for the tired on-call at 3am. Specific symptoms, fast triage, diagnostic tree, concrete mitigation steps with exact commands, escalation paths, verification. Test runbooks; keep them current; link from alerts and dashboards. Update after every incident. The shorter the runbook, the more likely it gets used; the more specific, the more likely it works.


Related reading



Keep learning. This article is part of the Advanced Quality Engineering path in the ShiftQuality Learning Center. Take quality from a team chore to an organizational property.



Part of the Engineering Templates & Playbooks guide — ShiftQuality's collection of ready-to-adapt templates.

bottom of page