Incident Management and Postmortems
A comprehensive guide to handling production incidents, establishing effective response processes, and learning from failures through blameless postmortems.
Overview
Incident management is the systematic approach to detecting, responding to, and recovering from service disruptions whilst minimising impact to users and the business. Effective incident management combines clear processes, well-defined roles, robust communication, and a culture of continuous improvement through structured postmortems.
flowchart TD
A[Alert Triggered] --> B{Severity Assessment}
B -->|SEV 1-2| C[Declare Incident]
B -->|SEV 3-4| D[Standard Response]
C --> E[Assemble War Room]
D --> F[Primary Responder]
E --> G[Assign Roles]
G --> H[IC: Incident Commander]
G --> I[TL: Technical Lead]
G --> J[CL: Comms Lead]
G --> K[Scribe]
H --> L[Investigation]
I --> L
J --> M[Stakeholder Updates]
K --> N[Document Timeline]
F --> L
L --> O{Root Cause Found?}
O -->|No| P[Continue Investigation]
P --> L
O -->|Yes| Q[Apply Mitigation]
Q --> R[Monitor Impact]
R --> S{Stable?}
S -->|No| T[Adjust Mitigation]
T --> R
S -->|Yes| U[Resolve Incident]
U --> V[Final Communication]
V --> W[Schedule Postmortem]
W --> X[Document Learnings]
X --> Y[Track Action Items]
Y --> Z[Implement Improvements]
style C fill:#ff6b6b
style E fill:#ff6b6b
style D fill:#ffd93d
style U fill:#6bcf7f
Triage and Severity Classification
Key Concepts
Proper triage ensures appropriate response urgency and resource allocation. The severity level determines escalation paths, stakeholder communication, and acceptable time-to-resolution.
Severity Levels
| Severity | Impact | Customer Experience | Response Time | Escalation | Examples |
|---|---|---|---|---|---|
| SEV 1 (Critical) | Complete service outage or critical security breach | All or majority of users cannot use core functionality | < 5 minutes | Page all hands, notify executives | Total site outage, data breach, payment processing down |
| SEV 2 (High) | Major feature degradation affecting significant users | Core functionality degraded or unavailable to subset of users | < 15 minutes | Page incident team, notify management | Authentication failures, checkout broken, database failover |
| SEV 3 (Medium) | Minor feature degradation with workaround available | Non-critical feature affected, degraded performance | < 1 hour | Notify on-call, work during business hours | Search slow, secondary feature unavailable, elevated error rates |
| SEV 4 (Low) | Minimal user impact, cosmetic issues | Minor inconvenience, edge case | Next business day | Standard work queue | UI typo, logging issues, minor performance degradation |
Triage Process
flowchart LR
A[Alert Received] --> B[Acknowledge Alert]
B --> C[Gather Initial Data]
C --> D{User Impact?}
D -->|Unknown| E[Check Metrics]
E --> F[Error Rates]
E --> G[Latency]
E --> H[Availability]
E --> I[Support Tickets]
F --> J{Impact Assessment}
G --> J
H --> J
I --> J
D -->|Known| J
J --> K{Scope}
K -->|100% users| L[SEV 1]
K -->|>25% users| M[SEV 2]
K -->|<25% users| N[SEV 3]
K -->|Minimal| O[SEV 4]
L --> P[Declare Incident]
M --> P
N --> Q[Standard Response]
O --> R[Create Ticket]
Severity Decision Matrix
User Impact Factors:
# SEV 1 Criteria (ANY of these)
sev_1:
- total_outage: true
- data_loss_risk: true
- security_breach: true
- payment_processing_down: true
- legal_regulatory_violation: true
# SEV 2 Criteria (ANY of these)
sev_2:
- core_feature_unavailable: true
- affected_users_percentage: ">25%"
- revenue_impact: "critical"
- slo_breach: ">50% error budget consumed in 1 hour"
# SEV 3 Criteria
sev_3:
- non_critical_feature_degraded: true
- affected_users_percentage: "<25%"
- workaround_available: true
- performance_degradation: "noticeable but functional"
# SEV 4 Criteria
sev_4:
- cosmetic_issue: true
- internal_tooling: true
- logging_monitoring_only: true
- edge_case: true
Triage Checklist
## Incident Triage Checklist
### Initial Assessment (0-5 minutes)
- [ ] Acknowledge alert in monitoring system
- [ ] Check service health dashboard
- [ ] Review recent deployments (last 24 hours)
- [ ] Check error rate and latency metrics
- [ ] Query support ticket volume
- [ ] Identify affected services/regions
### Impact Quantification
- [ ] Percentage of users affected: _____%
- [ ] Affected customer tier: [All/Premium/Free]
- [ ] Geographical scope: [Global/Regional/Single DC]
- [ ] Business impact: [Revenue/SLA/Reputation/None]
- [ ] Data integrity risk: [Yes/No]
### Severity Assignment
- [ ] Severity level determined: SEV-___
- [ ] Severity justification documented
- [ ] Escalation initiated if SEV 1-2
Auto-Triage with Automation
# Example automated triage logic
from dataclasses import dataclass
from typing import Optional
@dataclass
class IncidentMetrics:
error_rate: float # Percentage
latency_p99: float # Milliseconds
availability: float # Percentage
affected_users: int
total_users: int
has_data_loss_risk: bool = False
has_security_breach: bool = False
def determine_severity(metrics: IncidentMetrics) -> str:
"""Determine incident severity based on metrics."""
# SEV 1 conditions
if (
metrics.availability < 50 or
metrics.has_security_breach or
metrics.has_data_loss_risk or
metrics.error_rate > 50
):
return "SEV-1"
# SEV 2 conditions
affected_percentage = (metrics.affected_users / metrics.total_users) * 100
if (
metrics.availability < 99.0 or
metrics.error_rate > 10 or
affected_percentage > 25 or
metrics.latency_p99 > 5000 # 5 seconds
):
return "SEV-2"
# SEV 3 conditions
if (
metrics.availability < 99.9 or
metrics.error_rate > 1 or
affected_percentage > 5 or
metrics.latency_p99 > 2000 # 2 seconds
):
return "SEV-3"
return "SEV-4"
# Example usage
incident = IncidentMetrics(
error_rate=15.5,
latency_p99=3200,
availability=97.5,
affected_users=5000,
total_users=10000
)
severity = determine_severity(incident)
print(f"Recommended severity: {severity}")
# Output: Recommended severity: SEV-2
Communication Channels and Roles
Key Concepts
Effective incident response requires clear communication channels and well-defined roles to prevent confusion, reduce coordination overhead, and ensure all stakeholders receive timely updates.
Incident Roles and Responsibilities
graph TB
IC[Incident Commander]
TL[Technical Lead]
CL[Communications Lead]
SC[Scribe]
SME[Subject Matter Experts]
IC -->|Directs| TL
IC -->|Coordinates| CL
IC -->|Oversees| SC
IC -->|Requests| SME
TL -->|Investigates| TECH[Technical Systems]
TL -->|Implements| MIT[Mitigations]
CL -->|Updates| INT[Internal Stakeholders]
CL -->|Notifies| EXT[External Customers]
CL -->|Manages| STATUS[Status Page]
SC -->|Documents| TIME[Timeline]
SC -->|Records| DEC[Decisions]
SC -->|Tracks| ACT[Actions]
SME -->|Advises| TL
SME -->|Supports| TECH
style IC fill:#ff6b6b,color:#fff
style TL fill:#4ecdc4
style CL fill:#95e1d3
style SC fill:#ffd93d
style SME fill:#a8e6cf
| Role | Primary Responsibilities | Authority |
|---|---|---|
| Incident Commander (IC) | • Overall incident coordination<br>• Decision-making authority<br>• Resource allocation<br>• Handoff to new IC if needed<br>• Declare incident resolved | Full authority to make decisions, allocate resources, and override normal processes |
| Technical Lead (TL) | • Direct technical investigation<br>• Coordinate remediation efforts<br>• Evaluate mitigation options<br>• Communicate technical status to IC | Technical decisions, can request additional SMEs |
| Communications Lead (CL) | • Stakeholder updates (internal/external)<br>• Status page management<br>• Customer communication<br>• Executive briefings | Messaging and communication timing |
| Scribe | • Document incident timeline<br>• Record all decisions and actions<br>• Track who is doing what<br>• Capture metric snapshots | Read-only observation, factual recording |
| Subject Matter Experts (SMEs) | • Provide domain expertise<br>• Support investigation<br>• Implement fixes<br>• Validate hypotheses | Domain-specific technical decisions |
Communication Channels
Channel Structure:
# Slack/Teams channel naming convention
incident_channels:
primary: "#incident-{YYYY-MM-DD}-{service}-{number}"
example: "#incident-2026-01-15-payments-001"
# Channel purposes
purposes:
incident_primary:
- Technical discussion
- Decision making
- Coordination
- Real-time updates
incident_status:
- Customer-facing updates
- Stakeholder notifications
- Executive summaries
incident_postmortem:
- Post-incident discussion
- Action item tracking
- Learning documentation
# Video/Voice channels
war_room:
platform: "Zoom/Google Meet"
link_location: "Pinned in incident channel"
mandatory_for: ["SEV-1", "SEV-2"]
optional_for: ["SEV-3", "SEV-4"]
Channel Setup Template:
# Incident Channel Setup
## Channel: #incident-2026-01-15-auth-001
### Pinned Messages
**1. Incident Overview**
Severity: SEV-2 Service: Authentication Service Impact: Users unable to log in (estimated 35% of login attempts failing) Start Time: 2026-01-15 14:23 UTC Status: INVESTIGATING
War Room: https://meet.google.com/xxx-yyyy-zzz Status Page: https://status.example.com/incidents/1234 Dashboard: https://grafana.example.com/d/auth-service
**2. Roles**
Incident Commander: @alice Technical Lead: @bob Communications Lead: @charlie Scribe: @diana SMEs: @eve (auth), @frank (database)
**3. Quick Links**
Runbook: https://wiki.example.com/runbooks/auth-service Recent Deployments: https://github.com/org/auth/deployments Error Logs: https://logs.example.com/query/auth-errors Metrics: https://grafana.example.com/d/auth-metrics
### Channel Rules
- Technical discussion only - no social chat
- Use threads for detailed investigations
- Tag @here for critical updates only
- Scribe will record timeline - focus on solving
- Updates every 15 minutes minimum
Communication Cadence
| Severity | Update Frequency | Channels | Stakeholders |
|---|---|---|---|
| SEV 1 | Every 15 minutes | Incident channel, status page, executive brief, customer email | All hands, executives, customers |
| SEV 2 | Every 30 minutes | Incident channel, status page, internal Slack | Engineering leadership, affected teams, key customers |
| SEV 3 | Every 1-2 hours | Incident channel, team Slack | Team leadership, on-call |
| SEV 4 | Daily or at resolution | Ticket system | Assigned engineer |
Communication Templates
Incident Declaration (Internal):
@channel INCIDENT DECLARED - SEV-{X}
**Service:** {Service Name}
**Impact:** {Brief user-facing impact description}
**Affected:** {Number/percentage of users}
**Started:** {HH:MM UTC}
**Roles:**
- IC: @{name}
- TL: @{name}
- CL: @{name}
- Scribe: @{name}
**War Room:** {video call link}
**Dashboard:** {monitoring dashboard link}
**Status:** INVESTIGATING
Next update in 15 minutes or sooner if status changes.
Status Update (Internal):
**UPDATE - {HH:MM UTC}**
**Status:** {INVESTIGATING | IDENTIFIED | MITIGATING | MONITORING | RESOLVED}
**Summary:**
{What has happened since last update}
**Current Impact:**
{Latest user impact metrics}
**Actions Taken:**
- {Action 1}
- {Action 2}
**Next Steps:**
- {Planned action 1}
- {Planned action 2}
**ETA to Resolution:** {Time estimate or "Unknown"}
Next update: {HH:MM UTC}
Customer Communication (External):
Subject: [UPDATE] Service Disruption - {Service Name}
Dear Customers,
We are currently experiencing issues with {service name} that began at {HH:MM UTC}.
**Impact:**
{User-facing description of what isn't working}
**Current Status:**
We have identified the root cause and are implementing a fix. We expect service to be restored by {HH:MM UTC}.
**What you can do:**
{Workarounds if available, or "No action required on your part"}
We sincerely apologise for the inconvenience. We will send another update within {timeframe} or when the issue is resolved.
For real-time updates, please visit our status page: {status page URL}
Thank you for your patience.
The {Company} Team
Resolution Announcement:
@channel INCIDENT RESOLVED
**Service:** {Service Name}
**Duration:** {Start time} - {End time} ({total duration})
**Impact:** {Final impact summary}
**Resolution:**
{Brief description of how it was resolved}
**Root Cause:**
{High-level explanation - details in postmortem}
**Next Steps:**
- Postmortem scheduled for {date/time}
- Action items will be tracked in {ticket system}
- Full writeup will be shared by {date}
Thank you to everyone involved in the response.
Runbooks and Escalation Paths
Key Concepts
Runbooks provide step-by-step guidance for common incident scenarios, reducing cognitive load during high-stress situations. Escalation paths ensure appropriate expertise is engaged when needed.
Runbook Structure
# Runbook: {Service Name} - {Scenario}
**Metadata**
- Service: {service-name}
- Severity: {typical severity for this scenario}
- Owner: {team name}
- Last Updated: {YYYY-MM-DD}
- Slack Channel: #{team-channel}
## Symptoms
What you'll observe when this issue occurs:
- Error messages: {specific errors}
- Metrics: {affected metrics and thresholds}
- Alerts: {which alerts fire}
- User reports: {common complaints}
## Impact
- **User Impact:** {what users experience}
- **Affected Components:** {services/systems impacted}
- **Typical Severity:** SEV-{X}
## Quick Checks
Before starting investigation:
1. Check service health dashboard: {dashboard URL}
2. Verify recent deployments: {deployment history URL}
3. Check dependency status: {upstream/downstream services}
4. Review error rates: {metrics URL}
## Investigation Steps
### Step 1: Verify the Problem
```bash
# Check service health
kubectl get pods -n {namespace} -l app={service}
# Check recent logs
kubectl logs -n {namespace} -l app={service} --tail=100 --timestamps
# Query error metrics
# {Include specific metric queries}
Expected Output: {what you should see} If unexpected: {what to do}
Step 2: Identify Root Cause
Check these in order:
-
Database connectivity
# Test database connection {specific command} -
External dependencies
# Check API endpoints {specific commands} -
Resource constraints
# Check CPU/memory {specific commands}
Step 3: Mitigation Options
Choose appropriate mitigation based on root cause:
| Root Cause | Mitigation | Command | Risk |
|---|---|---|---|
| Bad deployment | Rollback to previous version | {rollback command} |
Low |
| Database overload | Scale read replicas | {scale command} |
Low |
| Memory leak | Restart pods rolling | {restart command} |
Medium |
| Dependency failure | Enable circuit breaker | {feature flag command} |
Low |
Step 4: Apply Mitigation
# Example: Rollback deployment
{detailed commands with explanation}
# Verify mitigation
{verification commands}
# Expected result
{what success looks like}
Step 5: Monitor Recovery
Watch these metrics for 30 minutes:
- Error rate: {should drop to < X%}
- Latency: {should return to < Xms}
- Throughput: {should recover to normal levels}
Escalation Criteria
Escalate to {team/person} if:
- [ ] Mitigation doesn't work within 15 minutes
- [ ] Impact is worsening
- [ ] Root cause is unclear after 30 minutes
- [ ] Multiple services affected
- [ ] Database integrity concerns
Escalation Contact: @{name} or #{channel}
Prevention
How to prevent this in future:
- {Prevention measure 1}
- {Prevention measure 2}
Related Runbooks
- {Link to related runbook 1}
- {Link to related runbook 2}
Additional Resources
- Architecture diagram: {URL}
- Service documentation: {URL}
- Historical incidents: {URL}
### Escalation Paths
```mermaid
flowchart TD
A[On-Call Engineer] -->|Cannot resolve in 15min| B{Severity?}
B -->|SEV 1| C[Page ALL HANDS]
B -->|SEV 2| D[Page Team Lead]
B -->|SEV 3-4| E[Consult SME]
C --> F[Incident Commander]
D --> F
E --> G{Resolved?}
G -->|No| D
G -->|Yes| H[Document & Close]
F --> I[Assemble War Room]
I --> J[Assign Roles]
J --> K{Progress in 30min?}
K -->|Yes| L[Continue]
K -->|No| M[Escalate to Manager]
L --> N{Resolved?}
N -->|No| K
N -->|Yes| O[Postmortem]
M --> P[Escalate to VP Engineering]
P --> Q{Still unresolved?}
Q -->|Yes| R[Executive Crisis Team]
Q -->|No| O
style C fill:#ff6b6b,color:#fff
style F fill:#ff6b6b,color:#fff
style R fill:#ff6b6b,color:#fff
Escalation Matrix:
| Condition | Primary Contact | Secondary Contact | Tertiary Contact |
|---|---|---|---|
| SEV 1 - Business Hours | Page all engineering | Notify VP Engineering | Notify CTO/CEO |
| SEV 1 - After Hours | Page primary on-call + team lead | Page secondary on-call + manager | Notify VP Engineering |
| SEV 2 - Business Hours | Team lead + relevant SMEs | Engineering manager | VP Engineering |
| SEV 2 - After Hours | Primary on-call | Secondary on-call | Team lead |
| Database Issues | Database SRE on-call | Database team lead | Database architect |
| Security Issues | Security on-call | Security team lead | CISO |
| Network Issues | Network SRE on-call | Network team lead | Infrastructure manager |
Escalation Contacts Configuration:
# Example PagerDuty escalation policy
escalation_policies:
- name: "Primary Engineering Escalation"
description: "Standard engineering incident escalation"
escalation_rules:
# Level 1: Primary on-call
- escalation_delay_minutes: 0
targets:
- type: schedule
id: primary_oncall_schedule
# Level 2: After 5 minutes, add secondary on-call
- escalation_delay_minutes: 5
targets:
- type: schedule
id: secondary_oncall_schedule
# Level 3: After 15 minutes, page team lead
- escalation_delay_minutes: 15
targets:
- type: user
id: team_lead_user_id
# Level 4: After 30 minutes, page engineering manager
- escalation_delay_minutes: 30
targets:
- type: user
id: engineering_manager_id
# Level 5: After 60 minutes, page VP Engineering
- escalation_delay_minutes: 60
targets:
- type: user
id: vp_engineering_id
- name: "SEV-1 Escalation"
description: "Immediate all-hands for critical incidents"
escalation_rules:
# All hands immediately
- escalation_delay_minutes: 0
targets:
- type: schedule
id: primary_oncall_schedule
- type: schedule
id: secondary_oncall_schedule
- type: user
id: team_lead_user_id
- type: user
id: engineering_manager_id
# Notify executives after 10 minutes if unresolved
- escalation_delay_minutes: 10
targets:
- type: user
id: vp_engineering_id
- type: user
id: cto_id
Runbook Discovery
Automated Runbook Suggestion:
# Example: Suggest relevant runbooks based on alert
import re
from typing import List, Dict
class RunbookMatcher:
def __init__(self, runbooks: List[Dict]):
self.runbooks = runbooks
def find_relevant_runbooks(
self,
alert_name: str,
service: str,
error_message: str
) -> List[Dict]:
"""Find runbooks matching the incident characteristics."""
relevant = []
for runbook in self.runbooks:
score = 0
# Match service
if runbook['service'].lower() == service.lower():
score += 3
# Match alert patterns
for symptom in runbook.get('symptoms', []):
if symptom.lower() in alert_name.lower():
score += 2
if symptom.lower() in error_message.lower():
score += 2
# Match keywords
for keyword in runbook.get('keywords', []):
if keyword.lower() in error_message.lower():
score += 1
if score > 0:
relevant.append({
'runbook': runbook,
'relevance_score': score
})
# Sort by relevance
relevant.sort(key=lambda x: x['relevance_score'], reverse=True)
return relevant
# Example usage
runbooks = [
{
'name': 'Database Connection Pool Exhaustion',
'service': 'payment-service',
'symptoms': ['high error rate', 'timeout', 'connection refused'],
'keywords': ['pool', 'connection', 'database'],
'url': 'https://wiki.example.com/runbooks/db-pool'
},
{
'name': 'Payment Gateway Timeout',
'service': 'payment-service',
'symptoms': ['payment failure', 'gateway timeout'],
'keywords': ['stripe', 'payment', 'gateway'],
'url': 'https://wiki.example.com/runbooks/payment-timeout'
}
]
matcher = RunbookMatcher(runbooks)
results = matcher.find_relevant_runbooks(
alert_name='PaymentServiceHighErrorRate',
service='payment-service',
error_message='connection timeout to database'
)
for result in results[:3]: # Top 3 matches
rb = result['runbook']
print(f"[Score: {result['relevance_score']}] {rb['name']}: {rb['url']}")
Blameless Postmortem Structure
Key Concepts
Blameless postmortems focus on systemic improvements rather than individual fault. The goal is to learn from failures, improve processes, and prevent recurrence without creating a culture of fear.
Blameless Culture Principles:
- Assume good intentions - People make the best decisions with information available at the time
- Focus on systems - Look for process/tooling gaps, not scapegoats
- Psychological safety - Encourage honest sharing without fear of punishment
- Learning over blaming - Treat failures as opportunities to improve
- Shared responsibility - Everyone contributes to reliability
flowchart LR
A[Incident Resolved] --> B[Schedule Postmortem]
B --> C[Gather Data]
C --> D[Timeline Reconstruction]
D --> E[Root Cause Analysis]
E --> F[Draft Postmortem]
F --> G[Team Review]
G --> H{Feedback?}
H -->|Yes| I[Incorporate Changes]
I --> F
H -->|No| J[Publish Postmortem]
J --> K[Present to Team]
K --> L[Extract Action Items]
L --> M[Assign Owners]
M --> N[Track Completion]
N --> O[Verify Implementation]
O --> P[Share Learnings]
style J fill:#6bcf7f
style P fill:#6bcf7f
Postmortem Timeline
| Timeframe | Activity |
|---|---|
| Within 2 hours | Scribe creates initial timeline from incident channel |
| Within 24 hours | Schedule postmortem meeting, identify participants |
| Within 48 hours | Draft postmortem document circulated for review |
| Within 5 days | Postmortem meeting conducted |
| Within 7 days | Final postmortem published, action items assigned |
| Within 14 days | First action item status review |
| Ongoing | Weekly action item progress tracking until complete |
Comprehensive Postmortem Template
# Postmortem: [Incident Title]
**Incident ID:** INC-{YYYY-MM-DD}-{number}
**Date:** {YYYY-MM-DD}
**Authors:** {Name 1}, {Name 2}
**Status:** [Draft | In Review | Final]
**Reviewers:** {Name 1}, {Name 2}
**Approvers:** {Engineering Manager}
---
## Executive Summary
[2-3 sentences covering: what broke, user impact, how it was fixed, key learnings]
**Example:**
On 15th January 2026, the payment processing service experienced a complete outage lasting 47 minutes affecting approximately 10,000 customers. The root cause was a database connection pool exhaustion triggered by a code change deployed earlier that day. The incident was resolved by rolling back the deployment. This postmortem identifies improvements to our testing, deployment, and monitoring practices.
---
## Impact Assessment
### User Impact
| Metric | Value |
|--------|-------|
| **Total Users Affected** | {number} users ({percentage}% of active users) |
| **Complete Service Loss** | {number} users |
| **Degraded Service** | {number} users |
| **Peak Error Rate** | {percentage}% |
| **Failed Transactions** | {number} |
| **User-Reported Issues** | {number} support tickets |
### Business Impact
| Metric | Value |
|--------|-------|
| **Revenue Lost** | £{amount} (estimated) |
| **SLA Breach** | {Yes/No} - {details if yes} |
| **Customer Credits Issued** | £{amount} |
| **Reputation Impact** | {Low/Medium/High} |
| **Media Coverage** | {None/Social/Press} |
### Service Impact
| Service | Status | Duration |
|---------|--------|----------|
| Payment Processing | Complete outage | 47 minutes |
| Order History | Degraded | 52 minutes |
| User Authentication | Normal | - |
### Error Budget Consumption
- **SLO Target:** 99.9% availability (43.2 minutes/month)
- **Budget Consumed:** 47 minutes
- **Remaining Budget:** -3.8 minutes (exhausted)
- **Action:** Feature freeze in effect until budget replenishes
---
## Timeline
All times in UTC. Key decision points highlighted in **bold**.
| Time | Event | Actor |
|------|-------|-------|
| 08:00 | Deployment of payment-service v3.2.0 began | CI/CD |
| 08:15 | Deployment completed successfully | CI/CD |
| 08:23 | First alerts: elevated error rate (5%) | Prometheus |
| 08:24 | On-call engineer acknowledged alert | @alice |
| 08:26 | Error rate spiking to 25% | Monitoring |
| 08:28 | **Incident declared as SEV-2** | @alice |
| 08:30 | War room established | @alice |
| 08:32 | Database team joined investigation | @bob (DBA) |
| 08:35 | Connection pool exhaustion identified | @bob |
| 08:37 | Recent deployment identified as suspect | @alice |
| 08:40 | **Decision: Rollback to v3.1.5** | @charlie (IC) |
| 08:42 | Rollback initiated | @alice |
| 08:47 | Rollback completed | CI/CD |
| 08:50 | Error rates dropping | Monitoring |
| 08:55 | All metrics returned to normal | Monitoring |
| 09:10 | **Incident resolved** (stable 15min) | @charlie (IC) |
| 09:25 | Status page updated: resolved | @diana (CL) |
| 10:00 | Code review identified connection leak | @eve |
**Total Duration:** 47 minutes (detection to resolution)
**Time to Detect:** 8 minutes (deployment to alert)
**Time to Mitigate:** 24 minutes (alert to mitigation started)
---
## Root Cause Analysis
### What Happened
The payment-service v3.2.0 deployment introduced a database connection leak in the payment validation code path. Under normal load, connections were acquired from the pool but not properly released after use. As the pool (configured with 100 max connections) gradually filled over 23 minutes, new payment requests began failing with "connection timeout" errors.
### Technical Details
**Problematic Code (v3.2.0):**
```python
# Bug: Connection not released in error path
def validate_payment(payment_id: str) -> bool:
conn = db_pool.get_connection() # Acquire connection
try:
result = conn.execute(
"SELECT status FROM payments WHERE id = ?",
(payment_id,)
)
return result.status == 'approved'
except DatabaseError:
# BUG: Connection not released on error
return False
finally:
conn.release() # Only reached on success path
Fix (reverted to v3.1.5):
# Corrected: Connection properly released in all paths
def validate_payment(payment_id: str) -> bool:
conn = db_pool.get_connection()
try:
result = conn.execute(
"SELECT status FROM payments WHERE id = ?",
(payment_id,)
)
return result.status == 'approved'
except DatabaseError:
return False
finally:
conn.release() # Always executed
Contributing Factors
-
Code Review Gap
- Resource management in error paths not systematically reviewed
- No checklist item for connection/resource lifecycle
- PR approved without load testing requirement
-
Testing Insufficiency
- Unit tests didn't cover error scenarios
- Integration tests used mocked database
- Load tests run for only 5 minutes (leak takes 20+ minutes to manifest)
-
Monitoring Gap
- No alerting on database connection pool utilisation
- Pool metrics collected but not monitored
- No pre-deployment canary analysis of pool metrics
-
Deployment Process
- No automated rollback on error rate spike
- Canary deployment only 5% traffic for 10 minutes
- Insufficient monitoring period before full rollout
Five Whys Analysis
-
Why did users experience payment failures?
- Because the payment service couldn't connect to the database
-
Why couldn't it connect to the database?
- Because the connection pool was exhausted (all 100 connections in use)
-
Why was the connection pool exhausted?
- Because connections were not being released back to the pool
-
Why weren't connections being released?
- Because the error handling code path didn't release the connection
-
Why did this code reach production?
- Because our testing doesn't validate resource cleanup and code review missed this pattern
Systemic Root Cause: Lack of automated verification for resource management in error paths
Detection and Response Evaluation
What Went Well ✓
- Fast Alert: Monitoring detected elevated errors within 8 minutes
- Rapid Acknowledgement: On-call responded in < 2 minutes
- Effective Escalation: Incident declared promptly, right severity assigned
- Good Communication: Status updates every 10 minutes, stakeholders informed
- Quick Mitigation: Rollback decision made decisively, executed efficiently
- Team Collaboration: Database expert engaged immediately, clear role assignment
- Documentation: Scribe maintained comprehensive timeline
What Could Be Improved ✗
- Detection Time: Issue existed 8 minutes before detection (should be < 5 min)
- Root Cause Identification: Took 12 minutes to identify connection pool issue
- Runbook Coverage: No existing runbook for connection pool exhaustion
- Automated Response: Manual rollback took 7 minutes (should be automated)
- Metric Visibility: Connection pool metrics not on main dashboard
- Canary Duration: 10-minute canary insufficient to detect slow leak
Response Metrics
| Metric | Target | Actual | Status |
|---|---|---|---|
| Time to Detect (TTD) | < 5 min | 8 min | ❌ Missed |
| Time to Acknowledge (TTA) | < 5 min | 1 min | ✅ Met |
| Time to Declare (TTI) | < 10 min | 4 min | ✅ Met |
| Time to Understand (TTU) | < 15 min | 12 min | ✅ Met |
| Time to Mitigate (TTM) | < 30 min | 24 min | ✅ Met |
| Time to Resolve (TTR) | < 60 min | 47 min | ✅ Met |
Action Items
All action items tracked in Jira project: RELIABILITY
| Priority | Action Item | Owner | Due Date | Status | Verification |
|---|---|---|---|---|---|
| P0 | Add connection pool utilisation alerting (>80% = warning, >95% = critical) | @bob | 2026-01-18 | ✅ Complete | Alert fires in staging |
| P0 | Implement automated rollback when error rate >10% for >5 minutes | @frank | 2026-01-22 | 🟡 In Progress | Tested in staging |
| P0 | Extend canary deployment monitoring from 10 to 30 minutes | @alice | 2026-01-19 | ✅ Complete | Updated in CI/CD config |
| P1 | Add connection pool metrics to main service dashboard | @grace | 2026-01-20 | ✅ Complete | Dashboard PR merged |
| P1 | Create runbook for database connection pool exhaustion | @bob | 2026-01-23 | 🟡 In Progress | Draft in review |
| P1 | Add resource management checklist to PR template | @charlie | 2026-01-21 | ✅ Complete | Template updated |
| P2 | Extend load test duration to 60 minutes minimum | @henry | 2026-01-25 | ⚪ Not Started | |
| P2 | Implement connection leak detection in test suite | @henry | 2026-01-30 | ⚪ Not Started | |
| P2 | Add static analysis rule for resource management patterns | @iris | 2026-02-05 | ⚪ Not Started | |
| P3 | Review all services for similar connection handling patterns | @team | 2026-02-15 | ⚪ Not Started |
Priority Definitions:
- P0: Critical, must fix immediately (complete within 1 week)
- P1: High priority, prevents recurrence (complete within 2 weeks)
- P2: Important, improves detection/response (complete within 1 month)
- P3: Nice to have, long-term improvements (complete within quarter)
Lessons Learned
Technical Lessons
-
Resource Management is Critical
- Learning: All acquired resources (connections, file handles, locks) must be released in all code paths, especially error handling
- Application: Add linting rules to detect missing resource cleanup, include in code review checklist
-
Load Tests Must Simulate Production Duration
- Learning: Short-duration load tests (<10 min) won't catch slow leaks
- Application: Implement "soak tests" running 60+ minutes with production-like traffic patterns
-
Connection Pool Metrics Are Essential
- Learning: Connection pool exhaustion is a common failure mode but often unmonitored
- Application: Make pool metrics a standard component of service dashboards and alerts
Process Lessons
-
Canary Deployments Need Adequate Duration
- Learning: 10-minute canaries are insufficient for detecting issues with slow onset
- Application: Extend canary phase to 30 minutes, monitor key resources, not just error rates
-
Automated Rollback Reduces MTTR
- Learning: Manual rollback took 7 minutes; automatic would be < 2 minutes
- Application: Implement automatic rollback triggers for error rate spikes
-
Runbooks Prevent Cognitive Overload
- Learning: Without a runbook, troubleshooting took longer during high-stress incident
- Application: Create runbooks for all common failure modes, review quarterly
Organisational Lessons
-
Error Budget Policy Works
- Learning: Error budget exhaustion triggered feature freeze, focusing team on reliability
- Application: Continue enforcing error budget policy, communicate broadly
-
Blameless Culture Enables Honesty
- Learning: Team member who wrote bug felt safe sharing details without fear
- Application: Continue reinforcing blameless postmortem culture in team communications
Supporting Information
Related Links
- Incident Ticket: INC-2026-01-15-001
- Incident Slack Channel: #incident-2026-01-15-payment-001
- War Room Recording: Google Meet Recording
- Pull Request with Bug: PR #2847
- Rollback PR: PR #2851
- Grafana Dashboard: Payment Service Overview
- Error Logs: Loki Query
Metrics and Graphs


Affected Versions
- Problematic Version: v3.2.0 (deployed 08:00 UTC, rolled back 08:42 UTC)
- Stable Version: v3.1.5 (current production version)
- Fix Version: v3.2.1 (scheduled for deployment 2026-01-17 after additional testing)
Postmortem Meeting Notes
Date: 2026-01-16 14:00 UTC Attendees: Alice (responder), Bob (DBA), Charlie (IC), Diana (CL), Eve (developer), Frank (SRE), Grace (monitoring)
Discussion Highlights:
- Team thanked for quick response and resolution
- No blame directed at developer who introduced bug - recognised as systemic issue
- Discussion of similar patterns in other services - to be audited
- Debate on automated rollback thresholds - decided on 10% error rate for 5 minutes
- Request for training on database connection management - training scheduled
- Suggestion to share learnings company-wide - postmortem to be presented at eng all-hands
Action Items Added During Meeting:
- P2: Schedule database connection management training (Owner: @bob, Due: 2026-01-30)
- P3: Present postmortem at engineering all-hands (Owner: @charlie, Due: 2026-01-19)
Appendix: Timeline Visualization
gantt
title Incident Timeline - Payment Service Outage
dateFormat HH:mm
axisFormat %H:%M
section Deployment
Deploy v3.2.0 :done, deploy, 08:00, 15m
section Incident
Error rate rising :crit, errors, 08:15, 28m
Complete outage :crit, outage, 08:26, 21m
section Response
Alert fires :active, alert, 08:23, 1m
On-call acknowledges :active, ack, 08:24, 2m
Investigation :active, invest, 08:26, 14m
Rollback execution :active, rollback, 08:40, 7m
section Recovery
Monitoring :done, monitor, 08:47, 23m
Incident resolved :milestone, 09:10, 0m
---
## Postmortem Review Checklist
Before publishing, verify:
- [ ] Executive summary clearly explains what happened to non-technical stakeholders
- [ ] Impact quantified with specific metrics (users, revenue, duration)
- [ ] Timeline is complete and accurate (verified against incident channel logs)
- [ ] Root cause is technical and specific, not vague
- [ ] Contributing factors identified (not just proximate cause)
- [ ] No blame language used anywhere in document
- [ ] "What went well" section included (not just problems)
- [ ] All action items have owners and due dates
- [ ] Action items are tracked in ticket system
- [ ] Lessons learned are actionable and specific
- [ ] Document reviewed by incident commander and technical lead
- [ ] Supporting links and evidence included
---
## Postmortem Distribution
**Internal Distribution:**
- Engineering team (Slack: #engineering)
- Product team (email)
- Customer success (email with customer-facing summary)
- Executive team (email with exec summary only)
- Company all-hands presentation (if SEV-1 or major learning)
**External Distribution (if applicable):**
- Public status page (customer-facing summary)
- Major customer accounts (personalised email)
- Blog post (for significant incidents affecting many customers)
**Retention:**
- Store in postmortem repository (wiki/Confluence)
- Tag by service, severity, failure type
- Index for searchability
- Review annually for pattern analysis
Action Item Tracking and Follow-Up
Key Concepts
Action items are worthless unless tracked to completion. Effective follow-up ensures learnings are translated into concrete improvements that prevent recurrence.
Action Item Lifecycle
stateDiagram-v2
[*] --> Identified: During postmortem
Identified --> Triaged: Priority assigned
Triaged --> Assigned: Owner and due date set
Assigned --> InProgress: Work started
InProgress --> InReview: Implementation complete
InReview --> Testing: Code review passed
Testing --> Verified: Testing complete
Verified --> Deployed: Rolled out to production
Deployed --> Validated: Verified in prod
Validated --> Closed: Documented & complete
Closed --> [*]
InProgress --> Blocked: Issue encountered
Blocked --> InProgress: Blocker resolved
Blocked --> Cancelled: No longer relevant
Cancelled --> [*]
Action Item Template
# Action Item Specification
action_item:
id: "PM-2026-015-001"
postmortem:
incident_id: "INC-2026-01-15-001"
incident_title: "Payment Service Database Connection Pool Exhaustion"
postmortem_url: "https://wiki.example.com/postmortems/2026-01-15-payment"
details:
title: "Add connection pool utilisation alerting"
description: |
Implement Prometheus alerts for database connection pool utilisation:
- Warning: >80% pool utilisation for 5 minutes
- Critical: >95% pool utilisation for 2 minutes
Alert should include:
- Current pool size and utilisation percentage
- Recent connection pool growth rate
- Link to runbook
- Link to service dashboard
priority: "P0" # P0, P1, P2, P3
category: "monitoring" # monitoring, testing, process, infrastructure, documentation
ownership:
owner: "@bob"
team: "database-sre"
reviewer: "@grace"
timeline:
created_date: "2026-01-16"
due_date: "2026-01-18"
started_date: "2026-01-16"
completed_date: "2026-01-17"
tracking:
jira_ticket: "REL-1234"
github_pr: "example/monitoring#567"
status: "complete" # not_started, in_progress, blocked, in_review, complete, cancelled
verification:
criteria: |
- Alert rules deployed to production Prometheus
- Alert fires in staging environment when pool >80%
- Runbook linked in alert annotations
- Alert routed to correct PagerDuty escalation
verified_by: "@grace"
verified_date: "2026-01-18"
verification_notes: "Tested in staging, alert fired correctly, routed to #database-alerts"
metrics:
estimated_mttr_improvement: "15 minutes"
estimated_recurrence_reduction: "90%"
Action Item Tracking Board
# Postmortem Action Items - Sprint View
## P0 - Critical (This Week)
| ID | Title | Owner | Due | Status | Blocker |
|----|-------|-------|-----|--------|---------|
| PM-2026-015-001 | Add connection pool alerting | @bob | Jan 18 | ✅ Complete | - |
| PM-2026-015-002 | Implement automated rollback | @frank | Jan 22 | 🟡 In Progress | Feature flag system |
| PM-2026-012-003 | Fix memory leak in auth service | @alice | Jan 19 | 🟡 In Progress | - |
## P1 - High (This Sprint)
| ID | Title | Owner | Due | Status | Blocker |
|----|-------|-------|-----|--------|---------|
| PM-2026-015-003 | Create connection pool runbook | @bob | Jan 23 | 🟡 In Progress | - |
| PM-2026-015-004 | Add pool metrics to dashboard | @grace | Jan 20 | ✅ Complete | - |
| PM-2026-013-001 | Implement circuit breaker for payment gateway | @eve | Jan 25 | ⚪ Not Started | - |
## P2 - Medium (This Month)
| ID | Title | Owner | Due | Status | Blocker |
|----|-------|-------|-----|--------|---------|
| PM-2026-015-005 | Extend load test duration | @henry | Jan 30 | ⚪ Not Started | - |
| PM-2026-014-002 | Add rate limiting to public API | @iris | Jan 28 | 🟡 In Progress | - |
## Blocked Items
| ID | Title | Owner | Blocker | Est. Resolution |
|----|-------|-------|---------|-----------------|
| PM-2026-015-002 | Automated rollback | @frank | Waiting for feature flag infrastructure | Jan 20 |
## Recently Completed
| ID | Title | Owner | Completed | Verification |
|----|-------|-------|-----------|--------------|
| PM-2026-015-001 | Connection pool alerting | @bob | Jan 17 | ✅ Verified in prod |
| PM-2026-015-004 | Pool metrics dashboard | @grace | Jan 19 | ✅ Verified in prod |
Follow-Up Cadence
Weekly Review:
## Action Items Review - Week of {Date}
**Attendees:** Engineering leadership, action item owners
### Agenda
1. **Review P0 Items** (5 min)
- All P0 items due this week
- Any blockers requiring escalation
2. **Review Blocked Items** (10 min)
- Current blockers
- Escalation needs
- Re-prioritisation if needed
3. **Review Completed Items** (5 min)
- Verification status
- Effectiveness assessment
- Lessons from implementation
4. **Update Priorities** (5 min)
- Adjust priorities based on new incidents
- Shift dates if capacity constraints
5. **Metrics Review** (5 min)
- Completion rate
- Overdue percentage
- Blocker trends
### Metrics This Week
- **Completion Rate:** 12/15 items (80%)
- **Overdue Items:** 2 (13%)
- **Average Time to Complete:**
- P0: 3.2 days (target: 7 days)
- P1: 9.1 days (target: 14 days)
- P2: 21.5 days (target: 30 days)
- **Blocked Items:** 3 (20%)
### Action Items from This Review
- [ ] Escalate feature flag blocker to VP Engineering (@frank)
- [ ] Allocate additional resource to overdue memory leak fix (@alice)
- [ ] Schedule training on circuit breaker patterns (@eve)
Monthly Retrospective:
## Monthly Postmortem Action Items Retrospective
**Period:** January 2026
### Completion Statistics
| Priority | Created | Completed | Cancelled | Outstanding | Completion Rate |
|----------|---------|-----------|-----------|-------------|-----------------|
| P0 | 8 | 8 | 0 | 0 | 100% |
| P1 | 15 | 12 | 1 | 2 | 86% |
| P2 | 22 | 14 | 3 | 5 | 70% |
| P3 | 18 | 6 | 2 | 10 | 40% |
| **Total** | **63** | **40** | **6** | **17** | **70%** |
### Impact Assessment
**Prevented Recurrence:**
- Connection pool exhaustion: Alert prevented 2 potential incidents
- Payment gateway timeout: Circuit breaker reduced impact by 90%
- Auth service memory leak: Issue resolved before customer impact
**Improved MTTR:**
- Average MTTR decreased from 45 to 32 minutes
- Automated rollback saved average 12 minutes per incident
**Reduced Toil:**
- Runbooks reduced investigation time by 40%
- Automated monitoring reduced false positive alerts by 60%
### Process Improvements
**What Worked Well:**
- Weekly review cadence kept items moving
- Clear ownership and due dates
- P0/P1 items prioritised effectively
- Good collaboration on blockers
**What Needs Improvement:**
- P2/P3 completion rate too low
- Some items created without clear verification criteria
- Blocked items languish too long
- Need better capacity planning for action items
### Action Items for Next Month
- [ ] Set target: 90% completion rate for P0/P1
- [ ] Require verification criteria before item creation
- [ ] Escalate blocked items after 3 days
- [ ] Reserve 20% engineering capacity for action items
Automated Action Item Tracking
# Example: Automated action item tracking and reminders
from dataclasses import dataclass
from datetime import date, timedelta
from typing import List, Optional
import enum
class Priority(enum.Enum):
P0 = 0
P1 = 1
P2 = 2
P3 = 3
class Status(enum.Enum):
NOT_STARTED = "not_started"
IN_PROGRESS = "in_progress"
BLOCKED = "blocked"
IN_REVIEW = "in_review"
COMPLETE = "complete"
CANCELLED = "cancelled"
@dataclass
class ActionItem:
id: str
title: str
owner: str
priority: Priority
due_date: date
status: Status
blocker: Optional[str] = None
jira_ticket: Optional[str] = None
class ActionItemTracker:
def __init__(self, items: List[ActionItem]):
self.items = items
def get_overdue_items(self) -> List[ActionItem]:
"""Return items past their due date."""
today = date.today()
return [
item for item in self.items
if item.due_date < today
and item.status not in [Status.COMPLETE, Status.CANCELLED]
]
def get_due_soon_items(self, days: int = 3) -> List[ActionItem]:
"""Return items due within N days."""
threshold = date.today() + timedelta(days=days)
return [
item for item in self.items
if item.due_date <= threshold
and item.status not in [Status.COMPLETE, Status.CANCELLED]
]
def get_blocked_items(self) -> List[ActionItem]:
"""Return currently blocked items."""
return [
item for item in self.items
if item.status == Status.BLOCKED
]
def send_reminders(self, slack_client):
"""Send automated reminders for action items."""
# Remind owners of overdue P0/P1 items
for item in self.get_overdue_items():
if item.priority in [Priority.P0, Priority.P1]:
slack_client.send_dm(
user=item.owner,
message=f"⚠️ **Overdue Action Item**\n\n"
f"**{item.title}** (ID: {item.id})\n"
f"Priority: {item.priority.name}\n"
f"Due: {item.due_date}\n"
f"Jira: {item.jira_ticket}\n\n"
f"Please update status or request help if blocked."
)
# Remind owners of items due soon
for item in self.get_due_soon_items():
slack_client.send_dm(
user=item.owner,
message=f"📅 **Action Item Due Soon**\n\n"
f"**{item.title}** (ID: {item.id})\n"
f"Priority: {item.priority.name}\n"
f"Due: {item.due_date}\n"
f"Jira: {item.jira_ticket}"
)
# Notify team channel of blocked items
blocked = self.get_blocked_items()
if blocked:
message = "🚧 **Blocked Action Items Requiring Attention**\n\n"
for item in blocked:
message += f"• **{item.title}** ({item.owner})\n"
message += f" Blocker: {item.blocker}\n\n"
slack_client.send_message(
channel="#reliability",
message=message
)
def get_completion_metrics(self) -> dict:
"""Calculate completion metrics."""
by_priority = {p: {"total": 0, "complete": 0} for p in Priority}
for item in self.items:
by_priority[item.priority]["total"] += 1
if item.status == Status.COMPLETE:
by_priority[item.priority]["complete"] += 1
return {
priority.name: {
"total": stats["total"],
"complete": stats["complete"],
"rate": (stats["complete"] / stats["total"] * 100) if stats["total"] > 0 else 0
}
for priority, stats in by_priority.items()
}
# Example usage
items = [
ActionItem(
id="PM-2026-015-001",
title="Add connection pool alerting",
owner="@bob",
priority=Priority.P0,
due_date=date(2026, 1, 18),
status=Status.COMPLETE
),
ActionItem(
id="PM-2026-015-002",
title="Implement automated rollback",
owner="@frank",
priority=Priority.P0,
due_date=date(2026, 1, 22),
status=Status.BLOCKED,
blocker="Waiting for feature flag infrastructure"
),
]
tracker = ActionItemTracker(items)
overdue = tracker.get_overdue_items()
metrics = tracker.get_completion_metrics()
print(f"Overdue P0/P1 items: {len(overdue)}")
print(f"P0 completion rate: {metrics['P0']['rate']:.1f}%")
Action Item Reporting Dashboard
# Postmortem Action Items Dashboard
## Overview
**Last Updated:** 2026-01-19 14:30 UTC
### Key Metrics
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| P0 Completion Rate | 100% | 100% | ✅ |
| P1 Completion Rate | 86% | 90% | 🟡 |
| P2 Completion Rate | 70% | 80% | ❌ |
| Average Days to Complete (P0) | 3.2 | < 7 | ✅ |
| Average Days to Complete (P1) | 9.1 | < 14 | ✅ |
| Overdue Items | 2 | 0 | ❌ |
| Blocked Items | 3 | < 5 | ✅ |
### Completion Trend
| Week | P0 | P1 | P2 | P3 | Total |
|---|---|---|---|---|---|
| W1 | 100% | 75% | 60% | 30% | 65% |
| W2 | 100% | 80% | 65% | 35% | 68% |
| W3 | 100% | 85% | 70% | 38% | 71% |
| W4 | 100% | 86% | 70% | 40% | 70% |
### By Category
| Category | Items | Complete | Rate |
|----------|-------|----------|------|
| Monitoring | 12 | 11 | 92% |
| Testing | 8 | 6 | 75% |
| Process | 15 | 10 | 67% |
| Infrastructure | 10 | 7 | 70% |
| Documentation | 18 | 6 | 33% |
### Top Contributors
| Person | Items Owned | Completed | Rate |
|--------|-------------|-----------|------|
| @bob | 8 | 7 | 88% |
| @alice | 6 | 5 | 83% |
| @frank | 5 | 3 | 60% |
| @grace | 4 | 4 | 100% |
Quick Reference
Severity Quick Decision Guide
| Question | SEV 1 | SEV 2 | SEV 3 | SEV 4 |
|---|---|---|---|---|
| Can users access core features? | No | Partially | Yes (slow) | Yes |
| Is money being lost? | Yes | Probably | Maybe | No |
| Is data at risk? | Yes | Potentially | No | No |
| What % of users affected? | >50% | 25-50% | 5-25% | <5% |
| Is there a workaround? | No | No | Yes | Yes |
| Response time? | <5 min | <15 min | <1 hr | Next day |
Incident Command Phrases
Useful phrases for Incident Commanders:
| Phrase | When to Use |
|---|---|
| "I need everyone on mute except who I call on" | War room getting chaotic |
| "Technical lead, what's your current hypothesis?" | Directing investigation |
| "Communications lead, send an update in 5 minutes" | Ensuring stakeholder communication |
| "Scribe, can you capture that decision?" | Documenting important choices |
| "Let's timebox this investigation to 10 minutes" | Preventing analysis paralysis |
| "I'm making the call to rollback" | Decisive mitigation decision |
| "We're going to try X. If it doesn't work in 15 minutes, we'll try Y" | Setting clear expectations |
| "This is resolved. Thank you everyone" | Clear incident closure |
Runbook Quick Checklist
Before declaring a runbook complete:
- [ ] Clear symptoms and impact defined
- [ ] Step-by-step investigation procedure
- [ ] Multiple mitigation options with trade-offs
- [ ] Verification steps for each mitigation
- [ ] Escalation criteria clearly stated
- [ ] All commands tested in staging
- [ ] Links to dashboards and logs
- [ ] Owner and last-updated date
- [ ] Peer reviewed by team
Postmortem Writing Tips
DO:
- ✅ Use passive voice: "The system failed" not "Bob broke it"
- ✅ Focus on systems: "Our testing didn't catch this" not "The engineer didn't test"
- ✅ Include what went well, not just problems
- ✅ Make action items specific and measurable
- ✅ Quantify impact with metrics
- ✅ Include timeline with timestamps
DON'T:
- ❌ Name individuals in relation to mistakes
- ❌ Use "should have" language (implies blame)
- ❌ Make vague action items like "improve testing"
- ❌ Skip the Five Whys analysis
- ❌ Forget to track action items to completion
- ❌ Rush the postmortem - quality over speed
Common Issues and Solutions
Issue: Severity Disagreement
Symptoms:
- On-call thinks SEV-3, manager thinks SEV-1
- Time wasted debating instead of responding
- Inconsistent severity application
Solutions:
- Use severity decision matrix - objective criteria
- Empower on-call to make initial call, adjust if needed
- When in doubt, declare higher severity (can always downgrade)
- Document reasoning for severity choice
- Review severity decisions in postmortem
Issue: Too Many Cooks in War Room
Symptoms:
- 20+ people in incident channel
- Multiple conflicting directions
- Technical leads confused about priorities
- Slow decision making
Solutions:
- Incident commander controls the room
- Mute all except active speakers
- Create observer-only channel for broader team
- Clearly assign roles - everyone else observers
- Use "raise hand" feature for questions
- IC explicitly calls on people to speak
Issue: Postmortems Don't Happen
Symptoms:
- Incidents resolved but no follow-up
- Same issues recurring
- No organisational learning
Solutions:
- Schedule postmortem within 24 hours (while fresh)
- Make postmortem completion a metric
- Tie postmortem completion to incident resolution
- Can't close incident ticket without postmortem
- Engineering leadership reviews postmortem completion rate
- Celebrate good postmortems publicly
Issue: Action Items Never Complete
Symptoms:
- Action items created but forgotten
- Low completion rate
- Recurring incidents from same root cause
Solutions:
- Track action items in main project management system
- Weekly review meeting with engineering leadership
- Reserve engineering capacity specifically for action items (20%)
- Make action item completion a performance factor
- Automate reminders for overdue items
- Publicly celebrate action item completion
- Escalate blocked items quickly
Issue: Blame Culture Preventing Honesty
Symptoms:
- Engineers hesitant to share mistakes
- Postmortems superficial, avoiding real issues
- "User error" or "bad luck" cited as root cause
- Fear of writing code or making changes
Solutions:
- Leadership models blameless behaviour
- Thank people for honesty in postmortems
- Never punish or embarrass for mistakes
- Focus all discussion on systems, not individuals
- Treat incidents as learning opportunities
- Publicly recognise people who own mistakes
- Make psychological safety a core value
Issue: Stakeholder Communication Gaps
Symptoms:
- Executives surprised by incidents
- Customers complaining about lack of updates
- Sales/support teams uninformed
Solutions:
- Assign dedicated communications lead for SEV-1/2
- Template-based updates reduce cognitive load
- Set update cadence and stick to it
- Use status page for customer communication
- Create internal vs external communication channels
- Brief executives immediately for SEV-1
- Provide support team with customer-facing messaging
Issue: Runbooks Out of Date
Symptoms:
- Commands in runbook don't work
- Runbook references old architecture
- Engineers don't trust runbooks
Solutions:
- Add "last updated" date to all runbooks
- Review runbooks quarterly
- Update runbook when it's used in real incident
- Make runbook accuracy an on-call responsibility
- Link runbooks to service ownership
- Test runbooks in chaos engineering exercises
- Archive outdated runbooks, don't leave them stale
Related Topics
The following topics would complement this Incident Management and Postmortems cheatsheet:
-
Observability Patterns - Comprehensive monitoring, alerting, and SLO practices that enable effective incident detection and response
-
Chaos Engineering - Proactive resilience testing including fault injection, game days, and building confidence in incident response procedures
-
Site Reliability Engineering (SRE) Practices - Broader SRE principles including error budgets, toil reduction, and production readiness reviews
-
On-Call Management - Deep dive into on-call schedules, escalation policies, alert design, and preventing on-call burnout
-
Prometheus & Grafana - Detailed coverage of metrics collection, alerting rules, dashboard design for incident detection
-
Distributed Tracing with OpenTelemetry - Instrumentation and trace analysis for debugging complex distributed system incidents