Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Incident Management and Postmortems

A comprehensive guide to handling production incidents, establishing effective response processes, and learning from failures through blameless postmortems.


Overview

Incident management is the systematic approach to detecting, responding to, and recovering from service disruptions whilst minimising impact to users and the business. Effective incident management combines clear processes, well-defined roles, robust communication, and a culture of continuous improvement through structured postmortems.

SEV 1-2SEV 3-4NoYesNoYesAlert TriggeredSeverity AssessmentDeclare IncidentStandard ResponseAssemble War RoomPrimary ResponderAssign RolesIC: IncidentCommanderTL: Technical LeadCL: Comms LeadScribeInvestigationStakeholder UpdatesDocument TimelineRoot Cause Found?ContinueInvestigationApply MitigationMonitor ImpactStable?Adjust MitigationResolve IncidentFinal CommunicationSchedule PostmortemDocument LearningsTrack Action ItemsImplementImprovementsSEV 1-2SEV 3-4NoYesNoYesAlert TriggeredSeverity AssessmentDeclare IncidentStandard ResponseAssemble War RoomPrimary ResponderAssign RolesIC: IncidentCommanderTL: Technical LeadCL: Comms LeadScribeInvestigationStakeholder UpdatesDocument TimelineRoot Cause Found?ContinueInvestigationApply MitigationMonitor ImpactStable?Adjust MitigationResolve IncidentFinal CommunicationSchedule PostmortemDocument LearningsTrack Action ItemsImplementImprovements

Triage and Severity Classification

Key Concepts

Proper triage ensures appropriate response urgency and resource allocation. The severity level determines escalation paths, stakeholder communication, and acceptable time-to-resolution.

Severity Levels

Severity Impact Customer Experience Response Time Escalation Examples
SEV 1 (Critical) Complete service outage or critical security breach All or majority of users cannot use core functionality < 5 minutes Page all hands, notify executives Total site outage, data breach, payment processing down
SEV 2 (High) Major feature degradation affecting significant users Core functionality degraded or unavailable to subset of users < 15 minutes Page incident team, notify management Authentication failures, checkout broken, database failover
SEV 3 (Medium) Minor feature degradation with workaround available Non-critical feature affected, degraded performance < 1 hour Notify on-call, work during business hours Search slow, secondary feature unavailable, elevated error rates
SEV 4 (Low) Minimal user impact, cosmetic issues Minor inconvenience, edge case Next business day Standard work queue UI typo, logging issues, minor performance degradation

Triage Process

UnknownKnown100% users>25% users&lt;25% usersMinimalAlert ReceivedAcknowledge AlertGather Initial DataUser Impact?Check MetricsError RatesLatencyAvailabilitySupport TicketsImpact AssessmentScopeSEV 1SEV 2SEV 3SEV 4Declare IncidentStandard ResponseCreate TicketUnknownKnown100% users>25% users&lt;25% usersMinimalAlert ReceivedAcknowledge AlertGather Initial DataUser Impact?Check MetricsError RatesLatencyAvailabilitySupport TicketsImpact AssessmentScopeSEV 1SEV 2SEV 3SEV 4Declare IncidentStandard ResponseCreate Ticket

Severity Decision Matrix

User Impact Factors:

# SEV 1 Criteria (ANY of these)
sev_1:
  - total_outage: true
  - data_loss_risk: true
  - security_breach: true
  - payment_processing_down: true
  - legal_regulatory_violation: true

# SEV 2 Criteria (ANY of these)
sev_2:
  - core_feature_unavailable: true
  - affected_users_percentage: ">25%"
  - revenue_impact: "critical"
  - slo_breach: ">50% error budget consumed in 1 hour"

# SEV 3 Criteria
sev_3:
  - non_critical_feature_degraded: true
  - affected_users_percentage: "<25%"
  - workaround_available: true
  - performance_degradation: "noticeable but functional"

# SEV 4 Criteria
sev_4:
  - cosmetic_issue: true
  - internal_tooling: true
  - logging_monitoring_only: true
  - edge_case: true

Triage Checklist

## Incident Triage Checklist

### Initial Assessment (0-5 minutes)
- [ ] Acknowledge alert in monitoring system
- [ ] Check service health dashboard
- [ ] Review recent deployments (last 24 hours)
- [ ] Check error rate and latency metrics
- [ ] Query support ticket volume
- [ ] Identify affected services/regions

### Impact Quantification
- [ ] Percentage of users affected: _____%
- [ ] Affected customer tier: [All/Premium/Free]
- [ ] Geographical scope: [Global/Regional/Single DC]
- [ ] Business impact: [Revenue/SLA/Reputation/None]
- [ ] Data integrity risk: [Yes/No]

### Severity Assignment
- [ ] Severity level determined: SEV-___
- [ ] Severity justification documented
- [ ] Escalation initiated if SEV 1-2

Auto-Triage with Automation

# Example automated triage logic
from dataclasses import dataclass
from typing import Optional

@dataclass
class IncidentMetrics:
    error_rate: float  # Percentage
    latency_p99: float  # Milliseconds
    availability: float  # Percentage
    affected_users: int
    total_users: int
    has_data_loss_risk: bool = False
    has_security_breach: bool = False

def determine_severity(metrics: IncidentMetrics) -> str:
    """Determine incident severity based on metrics."""

    # SEV 1 conditions
    if (
        metrics.availability < 50 or
        metrics.has_security_breach or
        metrics.has_data_loss_risk or
        metrics.error_rate > 50
    ):
        return "SEV-1"

    # SEV 2 conditions
    affected_percentage = (metrics.affected_users / metrics.total_users) * 100
    if (
        metrics.availability < 99.0 or
        metrics.error_rate > 10 or
        affected_percentage > 25 or
        metrics.latency_p99 > 5000  # 5 seconds
    ):
        return "SEV-2"

    # SEV 3 conditions
    if (
        metrics.availability < 99.9 or
        metrics.error_rate > 1 or
        affected_percentage > 5 or
        metrics.latency_p99 > 2000  # 2 seconds
    ):
        return "SEV-3"

    return "SEV-4"

# Example usage
incident = IncidentMetrics(
    error_rate=15.5,
    latency_p99=3200,
    availability=97.5,
    affected_users=5000,
    total_users=10000
)

severity = determine_severity(incident)
print(f"Recommended severity: {severity}")
# Output: Recommended severity: SEV-2

Communication Channels and Roles

Key Concepts

Effective incident response requires clear communication channels and well-defined roles to prevent confusion, reduce coordination overhead, and ensure all stakeholders receive timely updates.

Incident Roles and Responsibilities

DirectsCoordinatesOverseesRequestsInvestigatesImplementsUpdatesNotifiesManagesDocumentsRecordsTracksAdvisesSupportsIncident CommanderTechnical LeadCommunications LeadScribeSubject MatterExpertsTechnical SystemsMitigationsInternalStakeholdersExternal CustomersStatus PageTimelineDecisionsActionsDirectsCoordinatesOverseesRequestsInvestigatesImplementsUpdatesNotifiesManagesDocumentsRecordsTracksAdvisesSupportsIncident CommanderTechnical LeadCommunications LeadScribeSubject MatterExpertsTechnical SystemsMitigationsInternalStakeholdersExternal CustomersStatus PageTimelineDecisionsActions
Role Primary Responsibilities Authority
Incident Commander (IC) • Overall incident coordination<br>• Decision-making authority<br>• Resource allocation<br>• Handoff to new IC if needed<br>• Declare incident resolved Full authority to make decisions, allocate resources, and override normal processes
Technical Lead (TL) • Direct technical investigation<br>• Coordinate remediation efforts<br>• Evaluate mitigation options<br>• Communicate technical status to IC Technical decisions, can request additional SMEs
Communications Lead (CL) • Stakeholder updates (internal/external)<br>• Status page management<br>• Customer communication<br>• Executive briefings Messaging and communication timing
Scribe • Document incident timeline<br>• Record all decisions and actions<br>• Track who is doing what<br>• Capture metric snapshots Read-only observation, factual recording
Subject Matter Experts (SMEs) • Provide domain expertise<br>• Support investigation<br>• Implement fixes<br>• Validate hypotheses Domain-specific technical decisions

Communication Channels

Channel Structure:

# Slack/Teams channel naming convention
incident_channels:
  primary: "#incident-{YYYY-MM-DD}-{service}-{number}"
  example: "#incident-2026-01-15-payments-001"

  # Channel purposes
  purposes:
    incident_primary:
      - Technical discussion
      - Decision making
      - Coordination
      - Real-time updates

    incident_status:
      - Customer-facing updates
      - Stakeholder notifications
      - Executive summaries

    incident_postmortem:
      - Post-incident discussion
      - Action item tracking
      - Learning documentation

# Video/Voice channels
war_room:
  platform: "Zoom/Google Meet"
  link_location: "Pinned in incident channel"
  mandatory_for: ["SEV-1", "SEV-2"]
  optional_for: ["SEV-3", "SEV-4"]

Channel Setup Template:

# Incident Channel Setup

## Channel: #incident-2026-01-15-auth-001

### Pinned Messages

**1. Incident Overview**

Severity: SEV-2 Service: Authentication Service Impact: Users unable to log in (estimated 35% of login attempts failing) Start Time: 2026-01-15 14:23 UTC Status: INVESTIGATING

War Room: https://meet.google.com/xxx-yyyy-zzz Status Page: https://status.example.com/incidents/1234 Dashboard: https://grafana.example.com/d/auth-service

**2. Roles**

Incident Commander: @alice Technical Lead: @bob Communications Lead: @charlie Scribe: @diana SMEs: @eve (auth), @frank (database)

**3. Quick Links**

Runbook: https://wiki.example.com/runbooks/auth-service Recent Deployments: https://github.com/org/auth/deployments Error Logs: https://logs.example.com/query/auth-errors Metrics: https://grafana.example.com/d/auth-metrics

### Channel Rules
- Technical discussion only - no social chat
- Use threads for detailed investigations
- Tag @here for critical updates only
- Scribe will record timeline - focus on solving
- Updates every 15 minutes minimum

Communication Cadence

Severity Update Frequency Channels Stakeholders
SEV 1 Every 15 minutes Incident channel, status page, executive brief, customer email All hands, executives, customers
SEV 2 Every 30 minutes Incident channel, status page, internal Slack Engineering leadership, affected teams, key customers
SEV 3 Every 1-2 hours Incident channel, team Slack Team leadership, on-call
SEV 4 Daily or at resolution Ticket system Assigned engineer

Communication Templates

Incident Declaration (Internal):

@channel INCIDENT DECLARED - SEV-{X}

**Service:** {Service Name}
**Impact:** {Brief user-facing impact description}
**Affected:** {Number/percentage of users}
**Started:** {HH:MM UTC}

**Roles:**
- IC: @{name}
- TL: @{name}
- CL: @{name}
- Scribe: @{name}

**War Room:** {video call link}
**Dashboard:** {monitoring dashboard link}
**Status:** INVESTIGATING

Next update in 15 minutes or sooner if status changes.

Status Update (Internal):

**UPDATE - {HH:MM UTC}**

**Status:** {INVESTIGATING | IDENTIFIED | MITIGATING | MONITORING | RESOLVED}

**Summary:**
{What has happened since last update}

**Current Impact:**
{Latest user impact metrics}

**Actions Taken:**
- {Action 1}
- {Action 2}

**Next Steps:**
- {Planned action 1}
- {Planned action 2}

**ETA to Resolution:** {Time estimate or "Unknown"}

Next update: {HH:MM UTC}

Customer Communication (External):

Subject: [UPDATE] Service Disruption - {Service Name}

Dear Customers,

We are currently experiencing issues with {service name} that began at {HH:MM UTC}.

**Impact:**
{User-facing description of what isn't working}

**Current Status:**
We have identified the root cause and are implementing a fix. We expect service to be restored by {HH:MM UTC}.

**What you can do:**
{Workarounds if available, or "No action required on your part"}

We sincerely apologise for the inconvenience. We will send another update within {timeframe} or when the issue is resolved.

For real-time updates, please visit our status page: {status page URL}

Thank you for your patience.

The {Company} Team

Resolution Announcement:

@channel INCIDENT RESOLVED

**Service:** {Service Name}
**Duration:** {Start time} - {End time} ({total duration})
**Impact:** {Final impact summary}

**Resolution:**
{Brief description of how it was resolved}

**Root Cause:**
{High-level explanation - details in postmortem}

**Next Steps:**
- Postmortem scheduled for {date/time}
- Action items will be tracked in {ticket system}
- Full writeup will be shared by {date}

Thank you to everyone involved in the response.

Runbooks and Escalation Paths

Key Concepts

Runbooks provide step-by-step guidance for common incident scenarios, reducing cognitive load during high-stress situations. Escalation paths ensure appropriate expertise is engaged when needed.

Runbook Structure

# Runbook: {Service Name} - {Scenario}

**Metadata**
- Service: {service-name}
- Severity: {typical severity for this scenario}
- Owner: {team name}
- Last Updated: {YYYY-MM-DD}
- Slack Channel: #{team-channel}

## Symptoms

What you'll observe when this issue occurs:
- Error messages: {specific errors}
- Metrics: {affected metrics and thresholds}
- Alerts: {which alerts fire}
- User reports: {common complaints}

## Impact

- **User Impact:** {what users experience}
- **Affected Components:** {services/systems impacted}
- **Typical Severity:** SEV-{X}

## Quick Checks

Before starting investigation:

1. Check service health dashboard: {dashboard URL}
2. Verify recent deployments: {deployment history URL}
3. Check dependency status: {upstream/downstream services}
4. Review error rates: {metrics URL}

## Investigation Steps

### Step 1: Verify the Problem

```bash
# Check service health
kubectl get pods -n {namespace} -l app={service}

# Check recent logs
kubectl logs -n {namespace} -l app={service} --tail=100 --timestamps

# Query error metrics
# {Include specific metric queries}

Expected Output: {what you should see} If unexpected: {what to do}

Step 2: Identify Root Cause

Check these in order:

  1. Database connectivity

    # Test database connection
    {specific command}
    
  2. External dependencies

    # Check API endpoints
    {specific commands}
    
  3. Resource constraints

    # Check CPU/memory
    {specific commands}
    

Step 3: Mitigation Options

Choose appropriate mitigation based on root cause:

Root Cause Mitigation Command Risk
Bad deployment Rollback to previous version {rollback command} Low
Database overload Scale read replicas {scale command} Low
Memory leak Restart pods rolling {restart command} Medium
Dependency failure Enable circuit breaker {feature flag command} Low

Step 4: Apply Mitigation

# Example: Rollback deployment
{detailed commands with explanation}

# Verify mitigation
{verification commands}

# Expected result
{what success looks like}

Step 5: Monitor Recovery

Watch these metrics for 30 minutes:

  • Error rate: {should drop to < X%}
  • Latency: {should return to < Xms}
  • Throughput: {should recover to normal levels}

Escalation Criteria

Escalate to {team/person} if:

  • [ ] Mitigation doesn't work within 15 minutes
  • [ ] Impact is worsening
  • [ ] Root cause is unclear after 30 minutes
  • [ ] Multiple services affected
  • [ ] Database integrity concerns

Escalation Contact: @{name} or #{channel}

Prevention

How to prevent this in future:

  • {Prevention measure 1}
  • {Prevention measure 2}

Related Runbooks

  • {Link to related runbook 1}
  • {Link to related runbook 2}

Additional Resources

  • Architecture diagram: {URL}
  • Service documentation: {URL}
  • Historical incidents: {URL}
### Escalation Paths

```mermaid
flowchart TD
    A[On-Call Engineer] -->|Cannot resolve in 15min| B{Severity?}

    B -->|SEV 1| C[Page ALL HANDS]
    B -->|SEV 2| D[Page Team Lead]
    B -->|SEV 3-4| E[Consult SME]

    C --> F[Incident Commander]
    D --> F
    E --> G{Resolved?}

    G -->|No| D
    G -->|Yes| H[Document & Close]

    F --> I[Assemble War Room]
    I --> J[Assign Roles]

    J --> K{Progress in 30min?}
    K -->|Yes| L[Continue]
    K -->|No| M[Escalate to Manager]

    L --> N{Resolved?}
    N -->|No| K
    N -->|Yes| O[Postmortem]

    M --> P[Escalate to VP Engineering]
    P --> Q{Still unresolved?}
    Q -->|Yes| R[Executive Crisis Team]
    Q -->|No| O

    style C fill:#ff6b6b,color:#fff
    style F fill:#ff6b6b,color:#fff
    style R fill:#ff6b6b,color:#fff

Escalation Matrix:

Condition Primary Contact Secondary Contact Tertiary Contact
SEV 1 - Business Hours Page all engineering Notify VP Engineering Notify CTO/CEO
SEV 1 - After Hours Page primary on-call + team lead Page secondary on-call + manager Notify VP Engineering
SEV 2 - Business Hours Team lead + relevant SMEs Engineering manager VP Engineering
SEV 2 - After Hours Primary on-call Secondary on-call Team lead
Database Issues Database SRE on-call Database team lead Database architect
Security Issues Security on-call Security team lead CISO
Network Issues Network SRE on-call Network team lead Infrastructure manager

Escalation Contacts Configuration:

# Example PagerDuty escalation policy
escalation_policies:
  - name: "Primary Engineering Escalation"
    description: "Standard engineering incident escalation"

    escalation_rules:
      # Level 1: Primary on-call
      - escalation_delay_minutes: 0
        targets:
          - type: schedule
            id: primary_oncall_schedule

      # Level 2: After 5 minutes, add secondary on-call
      - escalation_delay_minutes: 5
        targets:
          - type: schedule
            id: secondary_oncall_schedule

      # Level 3: After 15 minutes, page team lead
      - escalation_delay_minutes: 15
        targets:
          - type: user
            id: team_lead_user_id

      # Level 4: After 30 minutes, page engineering manager
      - escalation_delay_minutes: 30
        targets:
          - type: user
            id: engineering_manager_id

      # Level 5: After 60 minutes, page VP Engineering
      - escalation_delay_minutes: 60
        targets:
          - type: user
            id: vp_engineering_id

  - name: "SEV-1 Escalation"
    description: "Immediate all-hands for critical incidents"

    escalation_rules:
      # All hands immediately
      - escalation_delay_minutes: 0
        targets:
          - type: schedule
            id: primary_oncall_schedule
          - type: schedule
            id: secondary_oncall_schedule
          - type: user
            id: team_lead_user_id
          - type: user
            id: engineering_manager_id

      # Notify executives after 10 minutes if unresolved
      - escalation_delay_minutes: 10
        targets:
          - type: user
            id: vp_engineering_id
          - type: user
            id: cto_id

Runbook Discovery

Automated Runbook Suggestion:

# Example: Suggest relevant runbooks based on alert
import re
from typing import List, Dict

class RunbookMatcher:
    def __init__(self, runbooks: List[Dict]):
        self.runbooks = runbooks

    def find_relevant_runbooks(
        self,
        alert_name: str,
        service: str,
        error_message: str
    ) -> List[Dict]:
        """Find runbooks matching the incident characteristics."""

        relevant = []

        for runbook in self.runbooks:
            score = 0

            # Match service
            if runbook['service'].lower() == service.lower():
                score += 3

            # Match alert patterns
            for symptom in runbook.get('symptoms', []):
                if symptom.lower() in alert_name.lower():
                    score += 2
                if symptom.lower() in error_message.lower():
                    score += 2

            # Match keywords
            for keyword in runbook.get('keywords', []):
                if keyword.lower() in error_message.lower():
                    score += 1

            if score > 0:
                relevant.append({
                    'runbook': runbook,
                    'relevance_score': score
                })

        # Sort by relevance
        relevant.sort(key=lambda x: x['relevance_score'], reverse=True)

        return relevant

# Example usage
runbooks = [
    {
        'name': 'Database Connection Pool Exhaustion',
        'service': 'payment-service',
        'symptoms': ['high error rate', 'timeout', 'connection refused'],
        'keywords': ['pool', 'connection', 'database'],
        'url': 'https://wiki.example.com/runbooks/db-pool'
    },
    {
        'name': 'Payment Gateway Timeout',
        'service': 'payment-service',
        'symptoms': ['payment failure', 'gateway timeout'],
        'keywords': ['stripe', 'payment', 'gateway'],
        'url': 'https://wiki.example.com/runbooks/payment-timeout'
    }
]

matcher = RunbookMatcher(runbooks)
results = matcher.find_relevant_runbooks(
    alert_name='PaymentServiceHighErrorRate',
    service='payment-service',
    error_message='connection timeout to database'
)

for result in results[:3]:  # Top 3 matches
    rb = result['runbook']
    print(f"[Score: {result['relevance_score']}] {rb['name']}: {rb['url']}")

Blameless Postmortem Structure

Key Concepts

Blameless postmortems focus on systemic improvements rather than individual fault. The goal is to learn from failures, improve processes, and prevent recurrence without creating a culture of fear.

Blameless Culture Principles:

  1. Assume good intentions - People make the best decisions with information available at the time
  2. Focus on systems - Look for process/tooling gaps, not scapegoats
  3. Psychological safety - Encourage honest sharing without fear of punishment
  4. Learning over blaming - Treat failures as opportunities to improve
  5. Shared responsibility - Everyone contributes to reliability
YesNoIncident ResolvedSchedule PostmortemGather DataTimelineReconstructionRoot Cause AnalysisDraft PostmortemTeam ReviewFeedback?Incorporate ChangesPublish PostmortemPresent to TeamExtract Action ItemsAssign OwnersTrack CompletionVerifyImplementationShare LearningsYesNoIncident ResolvedSchedule PostmortemGather DataTimelineReconstructionRoot Cause AnalysisDraft PostmortemTeam ReviewFeedback?Incorporate ChangesPublish PostmortemPresent to TeamExtract Action ItemsAssign OwnersTrack CompletionVerifyImplementationShare Learnings

Postmortem Timeline

Timeframe Activity
Within 2 hours Scribe creates initial timeline from incident channel
Within 24 hours Schedule postmortem meeting, identify participants
Within 48 hours Draft postmortem document circulated for review
Within 5 days Postmortem meeting conducted
Within 7 days Final postmortem published, action items assigned
Within 14 days First action item status review
Ongoing Weekly action item progress tracking until complete

Comprehensive Postmortem Template

# Postmortem: [Incident Title]

**Incident ID:** INC-{YYYY-MM-DD}-{number}
**Date:** {YYYY-MM-DD}
**Authors:** {Name 1}, {Name 2}
**Status:** [Draft | In Review | Final]
**Reviewers:** {Name 1}, {Name 2}
**Approvers:** {Engineering Manager}

---

## Executive Summary

[2-3 sentences covering: what broke, user impact, how it was fixed, key learnings]

**Example:**
On 15th January 2026, the payment processing service experienced a complete outage lasting 47 minutes affecting approximately 10,000 customers. The root cause was a database connection pool exhaustion triggered by a code change deployed earlier that day. The incident was resolved by rolling back the deployment. This postmortem identifies improvements to our testing, deployment, and monitoring practices.

---

## Impact Assessment

### User Impact

| Metric | Value |
|--------|-------|
| **Total Users Affected** | {number} users ({percentage}% of active users) |
| **Complete Service Loss** | {number} users |
| **Degraded Service** | {number} users |
| **Peak Error Rate** | {percentage}% |
| **Failed Transactions** | {number} |
| **User-Reported Issues** | {number} support tickets |

### Business Impact

| Metric | Value |
|--------|-------|
| **Revenue Lost** | £{amount} (estimated) |
| **SLA Breach** | {Yes/No} - {details if yes} |
| **Customer Credits Issued** | £{amount} |
| **Reputation Impact** | {Low/Medium/High} |
| **Media Coverage** | {None/Social/Press} |

### Service Impact

| Service | Status | Duration |
|---------|--------|----------|
| Payment Processing | Complete outage | 47 minutes |
| Order History | Degraded | 52 minutes |
| User Authentication | Normal | - |

### Error Budget Consumption

- **SLO Target:** 99.9% availability (43.2 minutes/month)
- **Budget Consumed:** 47 minutes
- **Remaining Budget:** -3.8 minutes (exhausted)
- **Action:** Feature freeze in effect until budget replenishes

---

## Timeline

All times in UTC. Key decision points highlighted in **bold**.

| Time | Event | Actor |
|------|-------|-------|
| 08:00 | Deployment of payment-service v3.2.0 began | CI/CD |
| 08:15 | Deployment completed successfully | CI/CD |
| 08:23 | First alerts: elevated error rate (5%) | Prometheus |
| 08:24 | On-call engineer acknowledged alert | @alice |
| 08:26 | Error rate spiking to 25% | Monitoring |
| 08:28 | **Incident declared as SEV-2** | @alice |
| 08:30 | War room established | @alice |
| 08:32 | Database team joined investigation | @bob (DBA) |
| 08:35 | Connection pool exhaustion identified | @bob |
| 08:37 | Recent deployment identified as suspect | @alice |
| 08:40 | **Decision: Rollback to v3.1.5** | @charlie (IC) |
| 08:42 | Rollback initiated | @alice |
| 08:47 | Rollback completed | CI/CD |
| 08:50 | Error rates dropping | Monitoring |
| 08:55 | All metrics returned to normal | Monitoring |
| 09:10 | **Incident resolved** (stable 15min) | @charlie (IC) |
| 09:25 | Status page updated: resolved | @diana (CL) |
| 10:00 | Code review identified connection leak | @eve |

**Total Duration:** 47 minutes (detection to resolution)
**Time to Detect:** 8 minutes (deployment to alert)
**Time to Mitigate:** 24 minutes (alert to mitigation started)

---

## Root Cause Analysis

### What Happened

The payment-service v3.2.0 deployment introduced a database connection leak in the payment validation code path. Under normal load, connections were acquired from the pool but not properly released after use. As the pool (configured with 100 max connections) gradually filled over 23 minutes, new payment requests began failing with "connection timeout" errors.

### Technical Details

**Problematic Code (v3.2.0):**

```python
# Bug: Connection not released in error path
def validate_payment(payment_id: str) -> bool:
    conn = db_pool.get_connection()  # Acquire connection
    try:
        result = conn.execute(
            "SELECT status FROM payments WHERE id = ?",
            (payment_id,)
        )
        return result.status == 'approved'
    except DatabaseError:
        # BUG: Connection not released on error
        return False
    finally:
        conn.release()  # Only reached on success path

Fix (reverted to v3.1.5):

# Corrected: Connection properly released in all paths
def validate_payment(payment_id: str) -> bool:
    conn = db_pool.get_connection()
    try:
        result = conn.execute(
            "SELECT status FROM payments WHERE id = ?",
            (payment_id,)
        )
        return result.status == 'approved'
    except DatabaseError:
        return False
    finally:
        conn.release()  # Always executed

Contributing Factors

  1. Code Review Gap

    • Resource management in error paths not systematically reviewed
    • No checklist item for connection/resource lifecycle
    • PR approved without load testing requirement
  2. Testing Insufficiency

    • Unit tests didn't cover error scenarios
    • Integration tests used mocked database
    • Load tests run for only 5 minutes (leak takes 20+ minutes to manifest)
  3. Monitoring Gap

    • No alerting on database connection pool utilisation
    • Pool metrics collected but not monitored
    • No pre-deployment canary analysis of pool metrics
  4. Deployment Process

    • No automated rollback on error rate spike
    • Canary deployment only 5% traffic for 10 minutes
    • Insufficient monitoring period before full rollout

Five Whys Analysis

  1. Why did users experience payment failures?

    • Because the payment service couldn't connect to the database
  2. Why couldn't it connect to the database?

    • Because the connection pool was exhausted (all 100 connections in use)
  3. Why was the connection pool exhausted?

    • Because connections were not being released back to the pool
  4. Why weren't connections being released?

    • Because the error handling code path didn't release the connection
  5. Why did this code reach production?

    • Because our testing doesn't validate resource cleanup and code review missed this pattern

Systemic Root Cause: Lack of automated verification for resource management in error paths


Detection and Response Evaluation

What Went Well ✓

  • Fast Alert: Monitoring detected elevated errors within 8 minutes
  • Rapid Acknowledgement: On-call responded in < 2 minutes
  • Effective Escalation: Incident declared promptly, right severity assigned
  • Good Communication: Status updates every 10 minutes, stakeholders informed
  • Quick Mitigation: Rollback decision made decisively, executed efficiently
  • Team Collaboration: Database expert engaged immediately, clear role assignment
  • Documentation: Scribe maintained comprehensive timeline

What Could Be Improved ✗

  • Detection Time: Issue existed 8 minutes before detection (should be < 5 min)
  • Root Cause Identification: Took 12 minutes to identify connection pool issue
  • Runbook Coverage: No existing runbook for connection pool exhaustion
  • Automated Response: Manual rollback took 7 minutes (should be automated)
  • Metric Visibility: Connection pool metrics not on main dashboard
  • Canary Duration: 10-minute canary insufficient to detect slow leak

Response Metrics

Metric Target Actual Status
Time to Detect (TTD) < 5 min 8 min ❌ Missed
Time to Acknowledge (TTA) < 5 min 1 min ✅ Met
Time to Declare (TTI) < 10 min 4 min ✅ Met
Time to Understand (TTU) < 15 min 12 min ✅ Met
Time to Mitigate (TTM) < 30 min 24 min ✅ Met
Time to Resolve (TTR) < 60 min 47 min ✅ Met

Action Items

All action items tracked in Jira project: RELIABILITY

Priority Action Item Owner Due Date Status Verification
P0 Add connection pool utilisation alerting (>80% = warning, >95% = critical) @bob 2026-01-18 ✅ Complete Alert fires in staging
P0 Implement automated rollback when error rate >10% for >5 minutes @frank 2026-01-22 🟡 In Progress Tested in staging
P0 Extend canary deployment monitoring from 10 to 30 minutes @alice 2026-01-19 ✅ Complete Updated in CI/CD config
P1 Add connection pool metrics to main service dashboard @grace 2026-01-20 ✅ Complete Dashboard PR merged
P1 Create runbook for database connection pool exhaustion @bob 2026-01-23 🟡 In Progress Draft in review
P1 Add resource management checklist to PR template @charlie 2026-01-21 ✅ Complete Template updated
P2 Extend load test duration to 60 minutes minimum @henry 2026-01-25 ⚪ Not Started
P2 Implement connection leak detection in test suite @henry 2026-01-30 ⚪ Not Started
P2 Add static analysis rule for resource management patterns @iris 2026-02-05 ⚪ Not Started
P3 Review all services for similar connection handling patterns @team 2026-02-15 ⚪ Not Started

Priority Definitions:

  • P0: Critical, must fix immediately (complete within 1 week)
  • P1: High priority, prevents recurrence (complete within 2 weeks)
  • P2: Important, improves detection/response (complete within 1 month)
  • P3: Nice to have, long-term improvements (complete within quarter)

Lessons Learned

Technical Lessons

  1. Resource Management is Critical

    • Learning: All acquired resources (connections, file handles, locks) must be released in all code paths, especially error handling
    • Application: Add linting rules to detect missing resource cleanup, include in code review checklist
  2. Load Tests Must Simulate Production Duration

    • Learning: Short-duration load tests (<10 min) won't catch slow leaks
    • Application: Implement "soak tests" running 60+ minutes with production-like traffic patterns
  3. Connection Pool Metrics Are Essential

    • Learning: Connection pool exhaustion is a common failure mode but often unmonitored
    • Application: Make pool metrics a standard component of service dashboards and alerts

Process Lessons

  1. Canary Deployments Need Adequate Duration

    • Learning: 10-minute canaries are insufficient for detecting issues with slow onset
    • Application: Extend canary phase to 30 minutes, monitor key resources, not just error rates
  2. Automated Rollback Reduces MTTR

    • Learning: Manual rollback took 7 minutes; automatic would be < 2 minutes
    • Application: Implement automatic rollback triggers for error rate spikes
  3. Runbooks Prevent Cognitive Overload

    • Learning: Without a runbook, troubleshooting took longer during high-stress incident
    • Application: Create runbooks for all common failure modes, review quarterly

Organisational Lessons

  1. Error Budget Policy Works

    • Learning: Error budget exhaustion triggered feature freeze, focusing team on reliability
    • Application: Continue enforcing error budget policy, communicate broadly
  2. Blameless Culture Enables Honesty

    • Learning: Team member who wrote bug felt safe sharing details without fear
    • Application: Continue reinforcing blameless postmortem culture in team communications

Supporting Information

Related Links

Metrics and Graphs

![Error Rate During Incident](https://grafana.example.com/render/d-solo/payment-overview/error-rate?from=1705308000&to=1705311600)

![Connection Pool Utilisation](https://grafana.example.com/render/d-solo/payment-overview/pool-utilisation?from=1705308000&to=1705311600)

Affected Versions

  • Problematic Version: v3.2.0 (deployed 08:00 UTC, rolled back 08:42 UTC)
  • Stable Version: v3.1.5 (current production version)
  • Fix Version: v3.2.1 (scheduled for deployment 2026-01-17 after additional testing)

Postmortem Meeting Notes

Date: 2026-01-16 14:00 UTC Attendees: Alice (responder), Bob (DBA), Charlie (IC), Diana (CL), Eve (developer), Frank (SRE), Grace (monitoring)

Discussion Highlights:

  • Team thanked for quick response and resolution
  • No blame directed at developer who introduced bug - recognised as systemic issue
  • Discussion of similar patterns in other services - to be audited
  • Debate on automated rollback thresholds - decided on 10% error rate for 5 minutes
  • Request for training on database connection management - training scheduled
  • Suggestion to share learnings company-wide - postmortem to be presented at eng all-hands

Action Items Added During Meeting:

  • P2: Schedule database connection management training (Owner: @bob, Due: 2026-01-30)
  • P3: Present postmortem at engineering all-hands (Owner: @charlie, Due: 2026-01-19)

Appendix: Timeline Visualization

08:0008:0508:1008:1508:2008:2508:3008:3508:4008:4508:5008:5509:0009:0509:10Deploy v3.2.0 Error rate rising Alert fires On-call acknowledges Complete outage Investigation Rollback execution Monitoring Incident resolved DeploymentIncidentResponseRecoveryIncident Timeline - Payment Service Outage08:0008:0508:1008:1508:2008:2508:3008:3508:4008:4508:5008:5509:0009:0509:10Deploy v3.2.0 Error rate rising Alert fires On-call acknowledges Complete outage Investigation Rollback execution Monitoring Incident resolved DeploymentIncidentResponseRecoveryIncident Timeline - Payment Service Outage
---

## Postmortem Review Checklist

Before publishing, verify:

- [ ] Executive summary clearly explains what happened to non-technical stakeholders
- [ ] Impact quantified with specific metrics (users, revenue, duration)
- [ ] Timeline is complete and accurate (verified against incident channel logs)
- [ ] Root cause is technical and specific, not vague
- [ ] Contributing factors identified (not just proximate cause)
- [ ] No blame language used anywhere in document
- [ ] "What went well" section included (not just problems)
- [ ] All action items have owners and due dates
- [ ] Action items are tracked in ticket system
- [ ] Lessons learned are actionable and specific
- [ ] Document reviewed by incident commander and technical lead
- [ ] Supporting links and evidence included

---

## Postmortem Distribution

**Internal Distribution:**
- Engineering team (Slack: #engineering)
- Product team (email)
- Customer success (email with customer-facing summary)
- Executive team (email with exec summary only)
- Company all-hands presentation (if SEV-1 or major learning)

**External Distribution (if applicable):**
- Public status page (customer-facing summary)
- Major customer accounts (personalised email)
- Blog post (for significant incidents affecting many customers)

**Retention:**
- Store in postmortem repository (wiki/Confluence)
- Tag by service, severity, failure type
- Index for searchability
- Review annually for pattern analysis

Action Item Tracking and Follow-Up

Key Concepts

Action items are worthless unless tracked to completion. Effective follow-up ensures learnings are translated into concrete improvements that prevent recurrence.

Action Item Lifecycle

During postmortemPriority assignedOwner and due date setWork startedImplementation completeCode review passedTesting completeRolled out to productionVerified in prodDocumented & completeIssue encounteredBlocker resolvedNo longer relevantIdentifiedTriagedAssignedInProgressInReviewTestingVerifiedDeployedValidatedClosedBlockedCancelledDuring postmortemPriority assignedOwner and due date setWork startedImplementation completeCode review passedTesting completeRolled out to productionVerified in prodDocumented & completeIssue encounteredBlocker resolvedNo longer relevantIdentifiedTriagedAssignedInProgressInReviewTestingVerifiedDeployedValidatedClosedBlockedCancelled

Action Item Template

# Action Item Specification
action_item:
  id: "PM-2026-015-001"

  postmortem:
    incident_id: "INC-2026-01-15-001"
    incident_title: "Payment Service Database Connection Pool Exhaustion"
    postmortem_url: "https://wiki.example.com/postmortems/2026-01-15-payment"

  details:
    title: "Add connection pool utilisation alerting"
    description: |
      Implement Prometheus alerts for database connection pool utilisation:
      - Warning: >80% pool utilisation for 5 minutes
      - Critical: >95% pool utilisation for 2 minutes

      Alert should include:
      - Current pool size and utilisation percentage
      - Recent connection pool growth rate
      - Link to runbook
      - Link to service dashboard

    priority: "P0"  # P0, P1, P2, P3
    category: "monitoring"  # monitoring, testing, process, infrastructure, documentation

  ownership:
    owner: "@bob"
    team: "database-sre"
    reviewer: "@grace"

  timeline:
    created_date: "2026-01-16"
    due_date: "2026-01-18"
    started_date: "2026-01-16"
    completed_date: "2026-01-17"

  tracking:
    jira_ticket: "REL-1234"
    github_pr: "example/monitoring#567"
    status: "complete"  # not_started, in_progress, blocked, in_review, complete, cancelled

  verification:
    criteria: |
      - Alert rules deployed to production Prometheus
      - Alert fires in staging environment when pool >80%
      - Runbook linked in alert annotations
      - Alert routed to correct PagerDuty escalation

    verified_by: "@grace"
    verified_date: "2026-01-18"
    verification_notes: "Tested in staging, alert fired correctly, routed to #database-alerts"

  metrics:
    estimated_mttr_improvement: "15 minutes"
    estimated_recurrence_reduction: "90%"

Action Item Tracking Board

# Postmortem Action Items - Sprint View

## P0 - Critical (This Week)

| ID | Title | Owner | Due | Status | Blocker |
|----|-------|-------|-----|--------|---------|
| PM-2026-015-001 | Add connection pool alerting | @bob | Jan 18 | ✅ Complete | - |
| PM-2026-015-002 | Implement automated rollback | @frank | Jan 22 | 🟡 In Progress | Feature flag system |
| PM-2026-012-003 | Fix memory leak in auth service | @alice | Jan 19 | 🟡 In Progress | - |

## P1 - High (This Sprint)

| ID | Title | Owner | Due | Status | Blocker |
|----|-------|-------|-----|--------|---------|
| PM-2026-015-003 | Create connection pool runbook | @bob | Jan 23 | 🟡 In Progress | - |
| PM-2026-015-004 | Add pool metrics to dashboard | @grace | Jan 20 | ✅ Complete | - |
| PM-2026-013-001 | Implement circuit breaker for payment gateway | @eve | Jan 25 | ⚪ Not Started | - |

## P2 - Medium (This Month)

| ID | Title | Owner | Due | Status | Blocker |
|----|-------|-------|-----|--------|---------|
| PM-2026-015-005 | Extend load test duration | @henry | Jan 30 | ⚪ Not Started | - |
| PM-2026-014-002 | Add rate limiting to public API | @iris | Jan 28 | 🟡 In Progress | - |

## Blocked Items

| ID | Title | Owner | Blocker | Est. Resolution |
|----|-------|-------|---------|-----------------|
| PM-2026-015-002 | Automated rollback | @frank | Waiting for feature flag infrastructure | Jan 20 |

## Recently Completed

| ID | Title | Owner | Completed | Verification |
|----|-------|-------|-----------|--------------|
| PM-2026-015-001 | Connection pool alerting | @bob | Jan 17 | ✅ Verified in prod |
| PM-2026-015-004 | Pool metrics dashboard | @grace | Jan 19 | ✅ Verified in prod |

Follow-Up Cadence

Weekly Review:

## Action Items Review - Week of {Date}

**Attendees:** Engineering leadership, action item owners

### Agenda

1. **Review P0 Items** (5 min)
   - All P0 items due this week
   - Any blockers requiring escalation

2. **Review Blocked Items** (10 min)
   - Current blockers
   - Escalation needs
   - Re-prioritisation if needed

3. **Review Completed Items** (5 min)
   - Verification status
   - Effectiveness assessment
   - Lessons from implementation

4. **Update Priorities** (5 min)
   - Adjust priorities based on new incidents
   - Shift dates if capacity constraints

5. **Metrics Review** (5 min)
   - Completion rate
   - Overdue percentage
   - Blocker trends

### Metrics This Week

- **Completion Rate:** 12/15 items (80%)
- **Overdue Items:** 2 (13%)
- **Average Time to Complete:**
  - P0: 3.2 days (target: 7 days)
  - P1: 9.1 days (target: 14 days)
  - P2: 21.5 days (target: 30 days)
- **Blocked Items:** 3 (20%)

### Action Items from This Review

- [ ] Escalate feature flag blocker to VP Engineering (@frank)
- [ ] Allocate additional resource to overdue memory leak fix (@alice)
- [ ] Schedule training on circuit breaker patterns (@eve)

Monthly Retrospective:

## Monthly Postmortem Action Items Retrospective

**Period:** January 2026

### Completion Statistics

| Priority | Created | Completed | Cancelled | Outstanding | Completion Rate |
|----------|---------|-----------|-----------|-------------|-----------------|
| P0 | 8 | 8 | 0 | 0 | 100% |
| P1 | 15 | 12 | 1 | 2 | 86% |
| P2 | 22 | 14 | 3 | 5 | 70% |
| P3 | 18 | 6 | 2 | 10 | 40% |
| **Total** | **63** | **40** | **6** | **17** | **70%** |

### Impact Assessment

**Prevented Recurrence:**
- Connection pool exhaustion: Alert prevented 2 potential incidents
- Payment gateway timeout: Circuit breaker reduced impact by 90%
- Auth service memory leak: Issue resolved before customer impact

**Improved MTTR:**
- Average MTTR decreased from 45 to 32 minutes
- Automated rollback saved average 12 minutes per incident

**Reduced Toil:**
- Runbooks reduced investigation time by 40%
- Automated monitoring reduced false positive alerts by 60%

### Process Improvements

**What Worked Well:**
- Weekly review cadence kept items moving
- Clear ownership and due dates
- P0/P1 items prioritised effectively
- Good collaboration on blockers

**What Needs Improvement:**
- P2/P3 completion rate too low
- Some items created without clear verification criteria
- Blocked items languish too long
- Need better capacity planning for action items

### Action Items for Next Month

- [ ] Set target: 90% completion rate for P0/P1
- [ ] Require verification criteria before item creation
- [ ] Escalate blocked items after 3 days
- [ ] Reserve 20% engineering capacity for action items

Automated Action Item Tracking

# Example: Automated action item tracking and reminders
from dataclasses import dataclass
from datetime import date, timedelta
from typing import List, Optional
import enum

class Priority(enum.Enum):
    P0 = 0
    P1 = 1
    P2 = 2
    P3 = 3

class Status(enum.Enum):
    NOT_STARTED = "not_started"
    IN_PROGRESS = "in_progress"
    BLOCKED = "blocked"
    IN_REVIEW = "in_review"
    COMPLETE = "complete"
    CANCELLED = "cancelled"

@dataclass
class ActionItem:
    id: str
    title: str
    owner: str
    priority: Priority
    due_date: date
    status: Status
    blocker: Optional[str] = None
    jira_ticket: Optional[str] = None

class ActionItemTracker:
    def __init__(self, items: List[ActionItem]):
        self.items = items

    def get_overdue_items(self) -> List[ActionItem]:
        """Return items past their due date."""
        today = date.today()
        return [
            item for item in self.items
            if item.due_date < today
            and item.status not in [Status.COMPLETE, Status.CANCELLED]
        ]

    def get_due_soon_items(self, days: int = 3) -> List[ActionItem]:
        """Return items due within N days."""
        threshold = date.today() + timedelta(days=days)
        return [
            item for item in self.items
            if item.due_date <= threshold
            and item.status not in [Status.COMPLETE, Status.CANCELLED]
        ]

    def get_blocked_items(self) -> List[ActionItem]:
        """Return currently blocked items."""
        return [
            item for item in self.items
            if item.status == Status.BLOCKED
        ]

    def send_reminders(self, slack_client):
        """Send automated reminders for action items."""

        # Remind owners of overdue P0/P1 items
        for item in self.get_overdue_items():
            if item.priority in [Priority.P0, Priority.P1]:
                slack_client.send_dm(
                    user=item.owner,
                    message=f"⚠️ **Overdue Action Item**\n\n"
                            f"**{item.title}** (ID: {item.id})\n"
                            f"Priority: {item.priority.name}\n"
                            f"Due: {item.due_date}\n"
                            f"Jira: {item.jira_ticket}\n\n"
                            f"Please update status or request help if blocked."
                )

        # Remind owners of items due soon
        for item in self.get_due_soon_items():
            slack_client.send_dm(
                user=item.owner,
                message=f"📅 **Action Item Due Soon**\n\n"
                        f"**{item.title}** (ID: {item.id})\n"
                        f"Priority: {item.priority.name}\n"
                        f"Due: {item.due_date}\n"
                        f"Jira: {item.jira_ticket}"
            )

        # Notify team channel of blocked items
        blocked = self.get_blocked_items()
        if blocked:
            message = "🚧 **Blocked Action Items Requiring Attention**\n\n"
            for item in blocked:
                message += f"• **{item.title}** ({item.owner})\n"
                message += f"  Blocker: {item.blocker}\n\n"

            slack_client.send_message(
                channel="#reliability",
                message=message
            )

    def get_completion_metrics(self) -> dict:
        """Calculate completion metrics."""
        by_priority = {p: {"total": 0, "complete": 0} for p in Priority}

        for item in self.items:
            by_priority[item.priority]["total"] += 1
            if item.status == Status.COMPLETE:
                by_priority[item.priority]["complete"] += 1

        return {
            priority.name: {
                "total": stats["total"],
                "complete": stats["complete"],
                "rate": (stats["complete"] / stats["total"] * 100) if stats["total"] > 0 else 0
            }
            for priority, stats in by_priority.items()
        }

# Example usage
items = [
    ActionItem(
        id="PM-2026-015-001",
        title="Add connection pool alerting",
        owner="@bob",
        priority=Priority.P0,
        due_date=date(2026, 1, 18),
        status=Status.COMPLETE
    ),
    ActionItem(
        id="PM-2026-015-002",
        title="Implement automated rollback",
        owner="@frank",
        priority=Priority.P0,
        due_date=date(2026, 1, 22),
        status=Status.BLOCKED,
        blocker="Waiting for feature flag infrastructure"
    ),
]

tracker = ActionItemTracker(items)
overdue = tracker.get_overdue_items()
metrics = tracker.get_completion_metrics()

print(f"Overdue P0/P1 items: {len(overdue)}")
print(f"P0 completion rate: {metrics['P0']['rate']:.1f}%")

Action Item Reporting Dashboard

# Postmortem Action Items Dashboard

## Overview

**Last Updated:** 2026-01-19 14:30 UTC

### Key Metrics

| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| P0 Completion Rate | 100% | 100% | ✅ |
| P1 Completion Rate | 86% | 90% | 🟡 |
| P2 Completion Rate | 70% | 80% | ❌ |
| Average Days to Complete (P0) | 3.2 | < 7 | ✅ |
| Average Days to Complete (P1) | 9.1 | < 14 | ✅ |
| Overdue Items | 2 | 0 | ❌ |
| Blocked Items | 3 | < 5 | ✅ |

### Completion Trend
Week P0 P1 P2 P3 Total
W1 100% 75% 60% 30% 65%
W2 100% 80% 65% 35% 68%
W3 100% 85% 70% 38% 71%
W4 100% 86% 70% 40% 70%
### By Category

| Category | Items | Complete | Rate |
|----------|-------|----------|------|
| Monitoring | 12 | 11 | 92% |
| Testing | 8 | 6 | 75% |
| Process | 15 | 10 | 67% |
| Infrastructure | 10 | 7 | 70% |
| Documentation | 18 | 6 | 33% |

### Top Contributors

| Person | Items Owned | Completed | Rate |
|--------|-------------|-----------|------|
| @bob | 8 | 7 | 88% |
| @alice | 6 | 5 | 83% |
| @frank | 5 | 3 | 60% |
| @grace | 4 | 4 | 100% |

Quick Reference

Severity Quick Decision Guide

Question SEV 1 SEV 2 SEV 3 SEV 4
Can users access core features? No Partially Yes (slow) Yes
Is money being lost? Yes Probably Maybe No
Is data at risk? Yes Potentially No No
What % of users affected? >50% 25-50% 5-25% <5%
Is there a workaround? No No Yes Yes
Response time? <5 min <15 min <1 hr Next day

Incident Command Phrases

Useful phrases for Incident Commanders:

Phrase When to Use
"I need everyone on mute except who I call on" War room getting chaotic
"Technical lead, what's your current hypothesis?" Directing investigation
"Communications lead, send an update in 5 minutes" Ensuring stakeholder communication
"Scribe, can you capture that decision?" Documenting important choices
"Let's timebox this investigation to 10 minutes" Preventing analysis paralysis
"I'm making the call to rollback" Decisive mitigation decision
"We're going to try X. If it doesn't work in 15 minutes, we'll try Y" Setting clear expectations
"This is resolved. Thank you everyone" Clear incident closure

Runbook Quick Checklist

Before declaring a runbook complete:

  • [ ] Clear symptoms and impact defined
  • [ ] Step-by-step investigation procedure
  • [ ] Multiple mitigation options with trade-offs
  • [ ] Verification steps for each mitigation
  • [ ] Escalation criteria clearly stated
  • [ ] All commands tested in staging
  • [ ] Links to dashboards and logs
  • [ ] Owner and last-updated date
  • [ ] Peer reviewed by team

Postmortem Writing Tips

DO:

  • ✅ Use passive voice: "The system failed" not "Bob broke it"
  • ✅ Focus on systems: "Our testing didn't catch this" not "The engineer didn't test"
  • ✅ Include what went well, not just problems
  • ✅ Make action items specific and measurable
  • ✅ Quantify impact with metrics
  • ✅ Include timeline with timestamps

DON'T:

  • ❌ Name individuals in relation to mistakes
  • ❌ Use "should have" language (implies blame)
  • ❌ Make vague action items like "improve testing"
  • ❌ Skip the Five Whys analysis
  • ❌ Forget to track action items to completion
  • ❌ Rush the postmortem - quality over speed

Common Issues and Solutions

Issue: Severity Disagreement

Symptoms:

  • On-call thinks SEV-3, manager thinks SEV-1
  • Time wasted debating instead of responding
  • Inconsistent severity application

Solutions:

  • Use severity decision matrix - objective criteria
  • Empower on-call to make initial call, adjust if needed
  • When in doubt, declare higher severity (can always downgrade)
  • Document reasoning for severity choice
  • Review severity decisions in postmortem

Issue: Too Many Cooks in War Room

Symptoms:

  • 20+ people in incident channel
  • Multiple conflicting directions
  • Technical leads confused about priorities
  • Slow decision making

Solutions:

  • Incident commander controls the room
  • Mute all except active speakers
  • Create observer-only channel for broader team
  • Clearly assign roles - everyone else observers
  • Use "raise hand" feature for questions
  • IC explicitly calls on people to speak

Issue: Postmortems Don't Happen

Symptoms:

  • Incidents resolved but no follow-up
  • Same issues recurring
  • No organisational learning

Solutions:

  • Schedule postmortem within 24 hours (while fresh)
  • Make postmortem completion a metric
  • Tie postmortem completion to incident resolution
  • Can't close incident ticket without postmortem
  • Engineering leadership reviews postmortem completion rate
  • Celebrate good postmortems publicly

Issue: Action Items Never Complete

Symptoms:

  • Action items created but forgotten
  • Low completion rate
  • Recurring incidents from same root cause

Solutions:

  • Track action items in main project management system
  • Weekly review meeting with engineering leadership
  • Reserve engineering capacity specifically for action items (20%)
  • Make action item completion a performance factor
  • Automate reminders for overdue items
  • Publicly celebrate action item completion
  • Escalate blocked items quickly

Issue: Blame Culture Preventing Honesty

Symptoms:

  • Engineers hesitant to share mistakes
  • Postmortems superficial, avoiding real issues
  • "User error" or "bad luck" cited as root cause
  • Fear of writing code or making changes

Solutions:

  • Leadership models blameless behaviour
  • Thank people for honesty in postmortems
  • Never punish or embarrass for mistakes
  • Focus all discussion on systems, not individuals
  • Treat incidents as learning opportunities
  • Publicly recognise people who own mistakes
  • Make psychological safety a core value

Issue: Stakeholder Communication Gaps

Symptoms:

  • Executives surprised by incidents
  • Customers complaining about lack of updates
  • Sales/support teams uninformed

Solutions:

  • Assign dedicated communications lead for SEV-1/2
  • Template-based updates reduce cognitive load
  • Set update cadence and stick to it
  • Use status page for customer communication
  • Create internal vs external communication channels
  • Brief executives immediately for SEV-1
  • Provide support team with customer-facing messaging

Issue: Runbooks Out of Date

Symptoms:

  • Commands in runbook don't work
  • Runbook references old architecture
  • Engineers don't trust runbooks

Solutions:

  • Add "last updated" date to all runbooks
  • Review runbooks quarterly
  • Update runbook when it's used in real incident
  • Make runbook accuracy an on-call responsibility
  • Link runbooks to service ownership
  • Test runbooks in chaos engineering exercises
  • Archive outdated runbooks, don't leave them stale

Related Topics

The following topics would complement this Incident Management and Postmortems cheatsheet:

  1. Observability Patterns - Comprehensive monitoring, alerting, and SLO practices that enable effective incident detection and response

  2. Chaos Engineering - Proactive resilience testing including fault injection, game days, and building confidence in incident response procedures

  3. Site Reliability Engineering (SRE) Practices - Broader SRE principles including error budgets, toil reduction, and production readiness reviews

  4. On-Call Management - Deep dive into on-call schedules, escalation policies, alert design, and preventing on-call burnout

  5. Prometheus & Grafana - Detailed coverage of metrics collection, alerting rules, dashboard design for incident detection

  6. Distributed Tracing with OpenTelemetry - Instrumentation and trace analysis for debugging complex distributed system incidents