Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Observability Patterns

A comprehensive guide to implementing observability practices for monitoring, alerting, and incident management in modern distributed systems.


Overview

Observability is the ability to understand the internal state of a system by examining its external outputs. Unlike traditional monitoring, observability enables you to ask arbitrary questions about your system's behaviour without deploying new instrumentation.

Feedback LoopActionAnalysis & VisualisationStorage & ProcessingData CollectionApplicationsMetricsLogsTracesPrometheus/InfluxDBElasticsearch/LokiJaeger/TempoGrafanaAlertmanagerPagerDuty/OpsgenieOn-Call EngineerIncident ResponsePostmortemImprovementsFeedback LoopActionAnalysis & VisualisationStorage & ProcessingData CollectionApplicationsMetricsLogsTracesPrometheus/InfluxDBElasticsearch/LokiJaeger/TempoGrafanaAlertmanagerPagerDuty/OpsgenieOn-Call EngineerIncident ResponsePostmortemImprovements

The Three Pillars of Observability

Key Concepts

Pillar Description Use Case Tools
Metrics Numerical measurements collected over time Performance trends, capacity planning Prometheus, InfluxDB, Datadog
Logs Timestamped records of discrete events Debugging, audit trails Elasticsearch, Loki, Splunk
Traces Request paths through distributed systems Latency analysis, dependency mapping Jaeger, Zipkin, Tempo

Metrics

Types of Metrics:

# Counter - only increases (e.g., requests served)
http_requests_total{method="GET", status="200"} 1234

# Gauge - can increase or decrease (e.g., temperature, queue size)
queue_depth{queue="orders"} 42

# Histogram - distribution of values (e.g., request duration)
http_request_duration_seconds_bucket{le="0.1"} 500
http_request_duration_seconds_bucket{le="0.5"} 800
http_request_duration_seconds_bucket{le="1.0"} 950

# Summary - similar to histogram with quantiles
http_request_duration_seconds{quantile="0.99"} 0.234

RED Method (for services):

  • Rate - requests per second
  • Errors - failed requests per second
  • Duration - time taken per request

USE Method (for resources):

  • Utilisation - percentage of resource busy
  • Saturation - amount of work queued
  • Errors - error count

Logs

Structured Logging Best Practices:

{
  "timestamp": "2024-01-15T10:30:00.000Z",
  "level": "ERROR",
  "service": "payment-service",
  "trace_id": "abc123def456",
  "span_id": "789xyz",
  "message": "Payment processing failed",
  "error": {
    "type": "PaymentGatewayError",
    "code": "INSUFFICIENT_FUNDS",
    "message": "Card declined"
  },
  "context": {
    "user_id": "user-12345",
    "order_id": "order-67890",
    "amount": 99.99,
    "currency": "GBP"
  }
}

Log Levels:

Level Use Case
TRACE Detailed debugging information
DEBUG Diagnostic information for developers
INFO General operational events
WARN Potentially harmful situations
ERROR Error events that allow continued operation
FATAL Severe errors causing application shutdown

Traces

OpenTelemetry Instrumentation Example:

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter

# Setup tracer
provider = TracerProvider()
processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317"))
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)

tracer = trace.get_tracer(__name__)

# Create spans
with tracer.start_as_current_span("process-order") as span:
    span.set_attribute("order.id", order_id)
    span.set_attribute("customer.id", customer_id)

    with tracer.start_as_current_span("validate-payment"):
        # Payment validation logic
        pass

    with tracer.start_as_current_span("update-inventory"):
        # Inventory update logic
        pass

Trace Context Propagation:

# W3C Trace Context Headers
traceparent: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01
tracestate: vendor1=value1,vendor2=value2

SLI/SLO/SLA Definitions

Key Concepts

Service LevelIndicatorWhat we measureService LevelObjectiveWhat we aim forService LevelAgreementWhat we promiseService LevelIndicatorWhat we measureService LevelObjectiveWhat we aim forService LevelAgreementWhat we promise
Term Definition Example
SLI Quantitative measure of service behaviour Request latency, error rate, throughput
SLO Target value or range for an SLI 99.9% of requests < 200ms
SLA Contract specifying consequences of missing SLO 99.5% availability or service credits

Common SLIs

Availability SLI:

# Availability = successful requests / total requests
sum(rate(http_requests_total{status!~"5.."}[30d]))
/
sum(rate(http_requests_total[30d]))

Latency SLI:

# Percentage of requests under 200ms
sum(rate(http_request_duration_seconds_bucket{le="0.2"}[30d]))
/
sum(rate(http_request_duration_seconds_count[30d]))

Throughput SLI:

# Requests processed per second
sum(rate(http_requests_total[5m]))

SLO Specification Example

apiVersion: sloth.slok.dev/v1
kind: PrometheusServiceLevel
metadata:
  name: payment-service
spec:
  service: "payment-service"
  labels:
    team: "payments"
    tier: "critical"
  slos:
    - name: "availability"
      objective: 99.9
      description: "Payment processing availability"
      sli:
        events:
          errorQuery: sum(rate(http_requests_total{service="payment", status=~"5.."}[{{.window}}]))
          totalQuery: sum(rate(http_requests_total{service="payment"}[{{.window}}]))
      alerting:
        name: PaymentAvailability
        labels:
          severity: critical
        pageAlert:
          labels:
            severity: page
        ticketAlert:
          labels:
            severity: ticket

Error Budgets

Key Concepts

An error budget is the maximum amount of time a service can fail without breaching its SLO over a given period.

Calculation:

Error Budget = 100% - SLO

For 99.9% SLO over 30 days:
Error Budget = 0.1%
Allowed Downtime = 30 days × 24 hours × 60 minutes × 0.001
                 = 43.2 minutes per month

Error Budget Consumption:

# Error budget remaining (as percentage)
1 - (
  (1 - (
    sum(rate(http_requests_total{status!~"5.."}[30d]))
    /
    sum(rate(http_requests_total[30d]))
  ))
  /
  (1 - 0.999)  # SLO target
)

Error Budget Policies

Budget Remaining Action
> 50% Normal development velocity
25-50% Reduce risky deployments
10-25% Focus on reliability work
< 10% Feature freeze, all hands on reliability
Exhausted No deployments until budget replenishes

Multi-Window Error Budget Alerts

# Fast-burn alert (high severity)
- alert: ErrorBudgetFastBurn
  expr: |
    (
      job:slo_errors_per_request:ratio_rate1h{job="api"} > (14.4 * 0.001)
      and
      job:slo_errors_per_request:ratio_rate5m{job="api"} > (14.4 * 0.001)
    )
  for: 2m
  labels:
    severity: critical
  annotations:
    summary: "High error rate consuming error budget rapidly"

# Slow-burn alert (lower severity)
- alert: ErrorBudgetSlowBurn
  expr: |
    (
      job:slo_errors_per_request:ratio_rate6h{job="api"} > (6 * 0.001)
      and
      job:slo_errors_per_request:ratio_rate30m{job="api"} > (6 * 0.001)
    )
  for: 15m
  labels:
    severity: warning
  annotations:
    summary: "Elevated error rate consuming error budget"

Alerting Best Practices

Key Concepts

Characteristics of Good Alerts:

  1. Actionable - Someone needs to do something
  2. Timely - Arrives with enough time to act
  3. Prioritised - Severity matches impact
  4. Contextual - Includes relevant information
  5. Unique - Avoids duplicate notifications

Alert Severity Levels

Severity Response Time Examples Notification
Critical/P1 Immediate (< 5 min) Service down, data loss risk Page on-call
High/P2 < 30 minutes Degraded performance, partial outage Page during business hours
Medium/P3 < 4 hours Non-critical component failure Slack/Teams notification
Low/P4 Next business day Informational, capacity planning Email/Ticket

Alert Configuration Example

groups:
  - name: api-alerts
    rules:
      # Good: Symptom-based, actionable alert
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
          /
          sum(rate(http_requests_total[5m]))
          > 0.01
        for: 5m
        labels:
          severity: critical
          team: platform
        annotations:
          summary: "API error rate above 1%"
          description: |
            Current error rate: {{ $value | humanizePercentage }}

            Runbook: https://wiki.example.com/runbooks/high-error-rate
            Dashboard: https://grafana.example.com/d/api-overview

      # Bad: Cause-based alert (avoid this)
      # - alert: HighCPU
      #   expr: cpu_usage > 80
      #   # CPU can be high without affecting users

      # Good: Alert on symptoms affecting users
      - alert: HighLatency
        expr: |
          histogram_quantile(0.99,
            sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
          ) > 1
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "P99 latency above 1 second"

Alert Fatigue Prevention

# Inhibition rules to prevent cascading alerts
inhibit_rules:
  # If critical is firing, suppress warning
  - source_match:
      severity: critical
    target_match:
      severity: warning
    equal: ['alertname', 'service']

  # If cluster is down, suppress individual node alerts
  - source_match:
      alertname: ClusterDown
    target_match_re:
      alertname: NodeDown|NodeDiskFull|NodeMemoryLow
    equal: ['cluster']

# Grouping to reduce notification volume
route:
  group_by: ['alertname', 'service', 'cluster']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h

On-Call Practices

Key Concepts

On-Call Responsibilities:

  1. Acknowledge alerts promptly
  2. Triage and assess impact
  3. Mitigate or escalate
  4. Document actions taken
  5. Hand off to next rotation

On-Call Schedule Best Practices

# Example PagerDuty schedule configuration
schedule:
  name: "Platform Team Primary"
  time_zone: "Europe/London"

  layers:
    - name: "Primary"
      rotation_virtual_start: "2024-01-01T09:00:00"
      rotation_turn_length_seconds: 604800  # 1 week
      users:
        - "P1234AB"  # Engineer 1
        - "P5678CD"  # Engineer 2
        - "P9012EF"  # Engineer 3
      restrictions:
        - type: "weekly_restriction"
          start_day_of_week: 1  # Monday
          start_time_of_day: "09:00:00"
          duration_seconds: 604800

escalation_policy:
  name: "Platform Escalation"
  escalation_rules:
    - escalation_delay_in_minutes: 5
      targets:
        - type: "schedule_reference"
          id: "primary-schedule-id"
    - escalation_delay_in_minutes: 15
      targets:
        - type: "user_reference"
          id: "team-lead-id"
    - escalation_delay_in_minutes: 30
      targets:
        - type: "user_reference"
          id: "manager-id"

On-Call Handoff Template

## On-Call Handoff - Week of [Date]

### Active Issues
- [ ] Issue #123: Database connection pool exhaustion (monitoring)
- [ ] Issue #456: Intermittent timeout to payment provider (ticket raised)

### Completed During Shift
- Resolved memory leak in order-service (PR #789 merged)
- Updated runbook for cache invalidation

### Upcoming Maintenance
- Tuesday 02:00 UTC: Database failover test
- Thursday 18:00 UTC: Kubernetes cluster upgrade

### Things to Watch
- Error budget at 35% for checkout-service
- New deployment of auth-service scheduled for Monday

### Notes for Next On-Call
- Payment provider having issues, be ready to enable backup
- New team member added to escalation path

On-Call Health Metrics

Metric Target Description
MTTA (Mean Time to Acknowledge) < 5 minutes Time from alert to acknowledgement
MTTR (Mean Time to Resolve) < 1 hour Time from alert to resolution
Interrupts per shift < 5 Number of pages per on-call shift
After-hours pages < 2/week Pages outside business hours
False positive rate < 10% Alerts that required no action

Incident Response Workflows

Key Concepts

Alert firesOn-call respondsImpact assessedSEV 3-4SEV 1-2War room assembledRoot cause foundFix appliedStable for 30minWithin 48hLearnings documentedDetectedAcknowledgedTriagedInvestigatingIncidentDeclaredMitigatingMonitoringResolvedPostmortemAlert firesOn-call respondsImpact assessedSEV 3-4SEV 1-2War room assembledRoot cause foundFix appliedStable for 30minWithin 48hLearnings documentedDetectedAcknowledgedTriagedInvestigatingIncidentDeclaredMitigatingMonitoringResolvedPostmortem

Incident Severity Definitions

Severity Impact Examples Response
SEV 1 Critical business impact Complete outage, data breach All hands, exec notification
SEV 2 Major feature degradation Core feature unavailable Incident commander + responders
SEV 3 Minor feature degradation Non-critical feature affected Primary responder
SEV 4 Minimal impact Cosmetic issue, edge case Normal work queue

Incident Roles

Role Responsibilities
Incident Commander (IC) Overall coordination, communication, decisions
Technical Lead Directs technical investigation and mitigation
Communications Lead Updates stakeholders, status page, customers
Scribe Documents timeline, actions, and decisions
Subject Matter Experts Provide domain expertise as needed

Incident Response Checklist

## Incident Response Checklist

### Initial Response (0-15 minutes)
- [ ] Acknowledge alert
- [ ] Join incident channel (#incident-[number])
- [ ] Assess initial impact and scope
- [ ] Assign severity level
- [ ] Notify incident commander if SEV 1-2

### Investigation (15-60 minutes)
- [ ] Establish timeline of events
- [ ] Review recent changes (deployments, configs)
- [ ] Check dashboards and logs
- [ ] Identify affected components
- [ ] Communicate status update every 15 minutes

### Mitigation
- [ ] Identify mitigation options
- [ ] Evaluate risks of each option
- [ ] Implement chosen mitigation
- [ ] Verify mitigation effectiveness
- [ ] Update status page

### Resolution
- [ ] Confirm service restored
- [ ] Monitor for recurrence (30 min minimum)
- [ ] Send final status update
- [ ] Schedule postmortem
- [ ] Update incident ticket with timeline

Communication Templates

Initial Incident Communication:

**Incident Declared - [Service Name]**

Severity: SEV-[1/2/3]
Impact: [Brief description of user impact]
Start Time: [HH:MM UTC]
Status: Investigating

Current Actions:
- [What team is doing]

Next Update: [Time] or sooner if status changes

Incident Channel: #incident-[number]
Status Page: https://status.example.com

Status Update:

**Update - [Service Name] Incident**

Time: [HH:MM UTC]
Status: [Investigating/Identified/Monitoring/Resolved]

Summary:
[What happened since last update]

Current Actions:
[What team is doing now]

ETA to Resolution: [Time estimate if known]

Next Update: [Time]

Postmortem Templates

Key Concepts

Postmortem Principles:

  • Blameless culture
  • Focus on systemic improvements
  • Share learnings broadly
  • Track action items to completion

Postmortem Template

# Postmortem: [Incident Title]

**Date:** [YYYY-MM-DD]
**Duration:** [Start time] - [End time] ([total duration])
**Severity:** SEV-[1/2/3/4]
**Authors:** [Names]
**Status:** [Draft/In Review/Final]

## Executive Summary

[2-3 sentences describing what happened, impact, and resolution]

## Impact

- **User Impact:** [Percentage of users affected, specific user journeys impacted]
- **Revenue Impact:** [Estimated financial impact if applicable]
- **Duration:** [Total time users experienced impact]
- **Support Tickets:** [Number of tickets received]

## Timeline

All times in UTC.

| Time | Event |
|------|-------|
| 09:15 | Deployment of auth-service v2.3.1 completed |
| 09:23 | First alerts fire for elevated error rates |
| 09:25 | On-call engineer acknowledges alert |
| 09:30 | Incident declared as SEV-2 |
| 09:35 | Identified recent deployment as potential cause |
| 09:40 | Initiated rollback to v2.3.0 |
| 09:45 | Rollback completed |
| 09:50 | Error rates returning to normal |
| 10:20 | Incident resolved after 30 min monitoring |

## Root Cause Analysis

### What Happened

[Detailed technical explanation of the failure chain]

### Contributing Factors

1. **Factor 1:** [Description]
   - Why it happened: [Explanation]

2. **Factor 2:** [Description]
   - Why it happened: [Explanation]

### Five Whys Analysis

1. Why did users see errors?
   - Because the auth service was returning 500 errors
2. Why was the auth service returning 500s?
   - Because it couldn't connect to the database
3. Why couldn't it connect to the database?
   - Because the connection pool was exhausted
4. Why was the connection pool exhausted?
   - Because the new code had a connection leak
5. Why wasn't this caught in testing?
   - Because load tests don't run long enough to detect slow leaks

## Detection

- **How was it detected:** [Automated alert / Customer report / Internal user]
- **Time to detect:** [Time from incident start to detection]
- **Detection gaps:** [What monitoring was missing or insufficient]

## Response

### What Went Well

- Alert fired within 2 minutes of impact starting
- On-call responded quickly
- Rollback procedure worked smoothly
- Communication was clear and timely

### What Could Be Improved

- Took 15 minutes to identify root cause
- Runbook was outdated for this scenario
- No automated rollback capability

## Action Items

| Priority | Action | Owner | Due Date | Status |
|----------|--------|-------|----------|--------|
| P1 | Add connection pool monitoring | @engineer1 | 2024-01-20 | In Progress |
| P1 | Implement automated rollback on error spike | @engineer2 | 2024-01-25 | Not Started |
| P2 | Extend load test duration to 4 hours | @engineer3 | 2024-01-22 | Not Started |
| P2 | Update auth-service runbook | @engineer1 | 2024-01-18 | Complete |
| P3 | Add connection leak detection to code review checklist | @tech-lead | 2024-01-30 | Not Started |

## Lessons Learned

1. **Lesson:** Connection pool metrics are critical for database-dependent services
   - **Application:** Add pool utilisation alerts for all services

2. **Lesson:** Load tests need to simulate production duration
   - **Application:** Create "soak test" category that runs for extended periods

3. **Lesson:** Automated rollback reduces MTTR significantly
   - **Application:** Prioritise feature flag integration for critical services

## Supporting Information

- **Incident Ticket:** INCIDENT-1234
- **Related PRs:** #567, #568
- **Dashboards:** [Link to relevant dashboards]
- **Logs:** [Link to log queries]
- **Slack Channel:** #incident-20240115-auth-outage

Action Item Tracking

# Example action item tracking in YAML
postmortem_actions:
  - id: PM-2024-001-01
    title: "Add connection pool monitoring"
    priority: P1
    owner: "@engineer1"
    due_date: "2024-01-20"
    status: "in_progress"
    verification: "Dashboard shows pool utilisation for all services"

  - id: PM-2024-001-02
    title: "Implement automated rollback"
    priority: P1
    owner: "@engineer2"
    due_date: "2024-01-25"
    status: "not_started"
    dependencies:
      - "Feature flag infrastructure"
    verification: "Rollback triggers automatically when error rate > 5%"

Quick Reference

Observability Metrics Formulas

Metric Formula Good Target
Availability (Total - Errors) / Total > 99.9%
Error Rate Errors / Total < 0.1%
P50 Latency 50th percentile response time < 100ms
P99 Latency 99th percentile response time < 500ms
Throughput Requests / Time Varies
MTTR Sum(resolution times) / Incidents < 1 hour
MTTA Sum(acknowledgement times) / Alerts < 5 minutes

Common PromQL Queries

# Request rate
sum(rate(http_requests_total[5m])) by (service)

# Error rate percentage
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))

# P99 latency
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))

# Availability over 30 days
avg_over_time(up{job="api"}[30d]) * 100

# Error budget burn rate
sum(rate(http_requests_total{status=~"5.."}[1h]))
  / sum(rate(http_requests_total[1h]))
  / (1 - 0.999)  # Divide by error budget (1 - SLO)

Incident Severity Quick Guide

Severity User Impact Response Example
SEV 1 Complete outage Page all + exec Site down
SEV 2 Major degradation Page team Checkout broken
SEV 3 Minor degradation Working hours Slow search
SEV 4 Minimal Next business day Typo in UI

Log Query Examples (Loki/LogQL)

# Errors in last hour
{service="api"} |= "error" | json | level="ERROR"

# Slow requests (> 1s)
{service="api"} | json | duration > 1000

# Requests by user
{service="api"} | json | user_id="user-123"

# Error rate by service
sum(rate({level="ERROR"}[5m])) by (service)

Common Issues and Solutions

Issue: Alert Fatigue

Symptoms:

  • On-call ignores alerts
  • High false positive rate
  • Multiple alerts for same issue

Solutions:

  • Implement alert inhibition rules
  • Add for duration to prevent flapping
  • Alert on symptoms, not causes
  • Review and prune alerts quarterly
  • Track alert actionability metrics

Issue: Missing Context in Alerts

Symptoms:

  • Alerts don't provide enough information
  • Engineers spend time gathering context
  • Runbooks not linked

Solutions:

annotations:
  summary: "Clear summary with current value: {{ $value }}"
  description: |
    **Impact:** Users cannot complete checkout
    **Dashboard:** https://grafana.example.com/d/checkout
    **Runbook:** https://wiki.example.com/runbooks/checkout-errors
    **Recent Changes:** https://github.com/org/repo/commits

Issue: Logs Not Correlated with Traces

Symptoms:

  • Cannot find logs for specific requests
  • Debugging distributed issues is difficult
  • Context lost between services

Solutions:

  • Inject trace_id and span_id into all logs
  • Use structured logging consistently
  • Configure log aggregator to index trace fields
  • Create dashboard links between trace and log views

Issue: SLO Not Reflecting User Experience

Symptoms:

  • SLOs green but users complaining
  • Metrics don't capture real issues
  • Business stakeholders distrustful of data

Solutions:

  • Define SLIs based on user journeys
  • Use synthetic monitoring for critical paths
  • Include client-side metrics
  • Validate SLOs with real user feedback
  • Regular SLO review meetings with stakeholders

Issue: Postmortems Not Leading to Improvements

Symptoms:

  • Same incidents recurring
  • Action items not completed
  • Learnings not shared

Solutions:

  • Track action items in issue tracker
  • Review action item completion weekly
  • Link postmortems to incident tickets
  • Share postmortems in team meetings
  • Create follow-up calendar reminders

Issue: On-Call Burnout

Symptoms:

  • High interrupt rate
  • Frequent night pages
  • Engineer turnover

Solutions:

  • Implement follow-the-sun rotations
  • Set target for < 2 after-hours pages/week
  • Invest in automation and self-healing
  • Allow on-call to focus solely on incidents
  • Compensate with time off or extra pay

Related Topics

The following topics would complement this Observability Patterns cheatsheet:

  1. Prometheus & Grafana - Deep dive into metrics collection, PromQL queries, and dashboard design for comprehensive monitoring setup

  2. Distributed Tracing with OpenTelemetry - Detailed coverage of instrumentation, context propagation, and trace analysis across microservices

  3. Chaos Engineering - Proactive resilience testing patterns including fault injection, game days, and steady-state hypothesis

  4. Site Reliability Engineering (SRE) Practices - Broader SRE principles including toil reduction, capacity planning, and production readiness

  5. Log Management with ELK/Loki - Log aggregation architecture, query languages, retention policies, and cost optimisation

  6. Kubernetes Monitoring - Container and orchestration-specific observability including resource metrics, cluster health, and workload monitoring