Observability Patterns
A comprehensive guide to implementing observability practices for monitoring, alerting, and incident management in modern distributed systems.
Overview
Observability is the ability to understand the internal state of a system by examining its external outputs. Unlike traditional monitoring, observability enables you to ask arbitrary questions about your system's behaviour without deploying new instrumentation.
graph TB
subgraph "Data Collection"
A[Applications] --> M[Metrics]
A --> L[Logs]
A --> T[Traces]
end
subgraph "Storage & Processing"
M --> PS[Prometheus/InfluxDB]
L --> ES[Elasticsearch/Loki]
T --> J[Jaeger/Tempo]
end
subgraph "Analysis & Visualisation"
PS --> G[Grafana]
ES --> G
J --> G
end
subgraph "Action"
G --> AL[Alertmanager]
AL --> PD[PagerDuty/Opsgenie]
PD --> OC[On-Call Engineer]
end
subgraph "Feedback Loop"
OC --> IR[Incident Response]
IR --> PM[Postmortem]
PM --> IMP[Improvements]
IMP --> A
end
The Three Pillars of Observability
Key Concepts
| Pillar | Description | Use Case | Tools |
|---|---|---|---|
| Metrics | Numerical measurements collected over time | Performance trends, capacity planning | Prometheus, InfluxDB, Datadog |
| Logs | Timestamped records of discrete events | Debugging, audit trails | Elasticsearch, Loki, Splunk |
| Traces | Request paths through distributed systems | Latency analysis, dependency mapping | Jaeger, Zipkin, Tempo |
Metrics
Types of Metrics:
# Counter - only increases (e.g., requests served)
http_requests_total{method="GET", status="200"} 1234
# Gauge - can increase or decrease (e.g., temperature, queue size)
queue_depth{queue="orders"} 42
# Histogram - distribution of values (e.g., request duration)
http_request_duration_seconds_bucket{le="0.1"} 500
http_request_duration_seconds_bucket{le="0.5"} 800
http_request_duration_seconds_bucket{le="1.0"} 950
# Summary - similar to histogram with quantiles
http_request_duration_seconds{quantile="0.99"} 0.234
RED Method (for services):
- Rate - requests per second
- Errors - failed requests per second
- Duration - time taken per request
USE Method (for resources):
- Utilisation - percentage of resource busy
- Saturation - amount of work queued
- Errors - error count
Logs
Structured Logging Best Practices:
{
"timestamp": "2024-01-15T10:30:00.000Z",
"level": "ERROR",
"service": "payment-service",
"trace_id": "abc123def456",
"span_id": "789xyz",
"message": "Payment processing failed",
"error": {
"type": "PaymentGatewayError",
"code": "INSUFFICIENT_FUNDS",
"message": "Card declined"
},
"context": {
"user_id": "user-12345",
"order_id": "order-67890",
"amount": 99.99,
"currency": "GBP"
}
}
Log Levels:
| Level | Use Case |
|---|---|
| TRACE | Detailed debugging information |
| DEBUG | Diagnostic information for developers |
| INFO | General operational events |
| WARN | Potentially harmful situations |
| ERROR | Error events that allow continued operation |
| FATAL | Severe errors causing application shutdown |
Traces
OpenTelemetry Instrumentation Example:
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
# Setup tracer
provider = TracerProvider()
processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317"))
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer(__name__)
# Create spans
with tracer.start_as_current_span("process-order") as span:
span.set_attribute("order.id", order_id)
span.set_attribute("customer.id", customer_id)
with tracer.start_as_current_span("validate-payment"):
# Payment validation logic
pass
with tracer.start_as_current_span("update-inventory"):
# Inventory update logic
pass
Trace Context Propagation:
# W3C Trace Context Headers
traceparent: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01
tracestate: vendor1=value1,vendor2=value2
SLI/SLO/SLA Definitions
Key Concepts
graph LR
SLI[Service Level Indicator<br/>What we measure] --> SLO[Service Level Objective<br/>What we aim for]
SLO --> SLA[Service Level Agreement<br/>What we promise]
style SLI fill:#e1f5fe
style SLO fill:#fff3e0
style SLA fill:#fce4ec
| Term | Definition | Example |
|---|---|---|
| SLI | Quantitative measure of service behaviour | Request latency, error rate, throughput |
| SLO | Target value or range for an SLI | 99.9% of requests < 200ms |
| SLA | Contract specifying consequences of missing SLO | 99.5% availability or service credits |
Common SLIs
Availability SLI:
# Availability = successful requests / total requests
sum(rate(http_requests_total{status!~"5.."}[30d]))
/
sum(rate(http_requests_total[30d]))
Latency SLI:
# Percentage of requests under 200ms
sum(rate(http_request_duration_seconds_bucket{le="0.2"}[30d]))
/
sum(rate(http_request_duration_seconds_count[30d]))
Throughput SLI:
# Requests processed per second
sum(rate(http_requests_total[5m]))
SLO Specification Example
apiVersion: sloth.slok.dev/v1
kind: PrometheusServiceLevel
metadata:
name: payment-service
spec:
service: "payment-service"
labels:
team: "payments"
tier: "critical"
slos:
- name: "availability"
objective: 99.9
description: "Payment processing availability"
sli:
events:
errorQuery: sum(rate(http_requests_total{service="payment", status=~"5.."}[{{.window}}]))
totalQuery: sum(rate(http_requests_total{service="payment"}[{{.window}}]))
alerting:
name: PaymentAvailability
labels:
severity: critical
pageAlert:
labels:
severity: page
ticketAlert:
labels:
severity: ticket
Error Budgets
Key Concepts
An error budget is the maximum amount of time a service can fail without breaching its SLO over a given period.
Calculation:
Error Budget = 100% - SLO
For 99.9% SLO over 30 days:
Error Budget = 0.1%
Allowed Downtime = 30 days × 24 hours × 60 minutes × 0.001
= 43.2 minutes per month
Error Budget Consumption:
# Error budget remaining (as percentage)
1 - (
(1 - (
sum(rate(http_requests_total{status!~"5.."}[30d]))
/
sum(rate(http_requests_total[30d]))
))
/
(1 - 0.999) # SLO target
)
Error Budget Policies
| Budget Remaining | Action |
|---|---|
| > 50% | Normal development velocity |
| 25-50% | Reduce risky deployments |
| 10-25% | Focus on reliability work |
| < 10% | Feature freeze, all hands on reliability |
| Exhausted | No deployments until budget replenishes |
Multi-Window Error Budget Alerts
# Fast-burn alert (high severity)
- alert: ErrorBudgetFastBurn
expr: |
(
job:slo_errors_per_request:ratio_rate1h{job="api"} > (14.4 * 0.001)
and
job:slo_errors_per_request:ratio_rate5m{job="api"} > (14.4 * 0.001)
)
for: 2m
labels:
severity: critical
annotations:
summary: "High error rate consuming error budget rapidly"
# Slow-burn alert (lower severity)
- alert: ErrorBudgetSlowBurn
expr: |
(
job:slo_errors_per_request:ratio_rate6h{job="api"} > (6 * 0.001)
and
job:slo_errors_per_request:ratio_rate30m{job="api"} > (6 * 0.001)
)
for: 15m
labels:
severity: warning
annotations:
summary: "Elevated error rate consuming error budget"
Alerting Best Practices
Key Concepts
Characteristics of Good Alerts:
- Actionable - Someone needs to do something
- Timely - Arrives with enough time to act
- Prioritised - Severity matches impact
- Contextual - Includes relevant information
- Unique - Avoids duplicate notifications
Alert Severity Levels
| Severity | Response Time | Examples | Notification |
|---|---|---|---|
| Critical/P1 | Immediate (< 5 min) | Service down, data loss risk | Page on-call |
| High/P2 | < 30 minutes | Degraded performance, partial outage | Page during business hours |
| Medium/P3 | < 4 hours | Non-critical component failure | Slack/Teams notification |
| Low/P4 | Next business day | Informational, capacity planning | Email/Ticket |
Alert Configuration Example
groups:
- name: api-alerts
rules:
# Good: Symptom-based, actionable alert
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
> 0.01
for: 5m
labels:
severity: critical
team: platform
annotations:
summary: "API error rate above 1%"
description: |
Current error rate: {{ $value | humanizePercentage }}
Runbook: https://wiki.example.com/runbooks/high-error-rate
Dashboard: https://grafana.example.com/d/api-overview
# Bad: Cause-based alert (avoid this)
# - alert: HighCPU
# expr: cpu_usage > 80
# # CPU can be high without affecting users
# Good: Alert on symptoms affecting users
- alert: HighLatency
expr: |
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
) > 1
for: 10m
labels:
severity: warning
annotations:
summary: "P99 latency above 1 second"
Alert Fatigue Prevention
# Inhibition rules to prevent cascading alerts
inhibit_rules:
# If critical is firing, suppress warning
- source_match:
severity: critical
target_match:
severity: warning
equal: ['alertname', 'service']
# If cluster is down, suppress individual node alerts
- source_match:
alertname: ClusterDown
target_match_re:
alertname: NodeDown|NodeDiskFull|NodeMemoryLow
equal: ['cluster']
# Grouping to reduce notification volume
route:
group_by: ['alertname', 'service', 'cluster']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
On-Call Practices
Key Concepts
On-Call Responsibilities:
- Acknowledge alerts promptly
- Triage and assess impact
- Mitigate or escalate
- Document actions taken
- Hand off to next rotation
On-Call Schedule Best Practices
# Example PagerDuty schedule configuration
schedule:
name: "Platform Team Primary"
time_zone: "Europe/London"
layers:
- name: "Primary"
rotation_virtual_start: "2024-01-01T09:00:00"
rotation_turn_length_seconds: 604800 # 1 week
users:
- "P1234AB" # Engineer 1
- "P5678CD" # Engineer 2
- "P9012EF" # Engineer 3
restrictions:
- type: "weekly_restriction"
start_day_of_week: 1 # Monday
start_time_of_day: "09:00:00"
duration_seconds: 604800
escalation_policy:
name: "Platform Escalation"
escalation_rules:
- escalation_delay_in_minutes: 5
targets:
- type: "schedule_reference"
id: "primary-schedule-id"
- escalation_delay_in_minutes: 15
targets:
- type: "user_reference"
id: "team-lead-id"
- escalation_delay_in_minutes: 30
targets:
- type: "user_reference"
id: "manager-id"
On-Call Handoff Template
## On-Call Handoff - Week of [Date]
### Active Issues
- [ ] Issue #123: Database connection pool exhaustion (monitoring)
- [ ] Issue #456: Intermittent timeout to payment provider (ticket raised)
### Completed During Shift
- Resolved memory leak in order-service (PR #789 merged)
- Updated runbook for cache invalidation
### Upcoming Maintenance
- Tuesday 02:00 UTC: Database failover test
- Thursday 18:00 UTC: Kubernetes cluster upgrade
### Things to Watch
- Error budget at 35% for checkout-service
- New deployment of auth-service scheduled for Monday
### Notes for Next On-Call
- Payment provider having issues, be ready to enable backup
- New team member added to escalation path
On-Call Health Metrics
| Metric | Target | Description |
|---|---|---|
| MTTA (Mean Time to Acknowledge) | < 5 minutes | Time from alert to acknowledgement |
| MTTR (Mean Time to Resolve) | < 1 hour | Time from alert to resolution |
| Interrupts per shift | < 5 | Number of pages per on-call shift |
| After-hours pages | < 2/week | Pages outside business hours |
| False positive rate | < 10% | Alerts that required no action |
Incident Response Workflows
Key Concepts
stateDiagram-v2
[*] --> Detected: Alert fires
Detected --> Acknowledged: On-call responds
Acknowledged --> Triaged: Impact assessed
Triaged --> Investigating: SEV 3-4
Triaged --> IncidentDeclared: SEV 1-2
IncidentDeclared --> Mitigating: War room assembled
Investigating --> Mitigating: Root cause found
Mitigating --> Monitoring: Fix applied
Monitoring --> Resolved: Stable for 30min
Resolved --> Postmortem: Within 48h
Postmortem --> [*]: Learnings documented
Incident Severity Definitions
| Severity | Impact | Examples | Response |
|---|---|---|---|
| SEV 1 | Critical business impact | Complete outage, data breach | All hands, exec notification |
| SEV 2 | Major feature degradation | Core feature unavailable | Incident commander + responders |
| SEV 3 | Minor feature degradation | Non-critical feature affected | Primary responder |
| SEV 4 | Minimal impact | Cosmetic issue, edge case | Normal work queue |
Incident Roles
| Role | Responsibilities |
|---|---|
| Incident Commander (IC) | Overall coordination, communication, decisions |
| Technical Lead | Directs technical investigation and mitigation |
| Communications Lead | Updates stakeholders, status page, customers |
| Scribe | Documents timeline, actions, and decisions |
| Subject Matter Experts | Provide domain expertise as needed |
Incident Response Checklist
## Incident Response Checklist
### Initial Response (0-15 minutes)
- [ ] Acknowledge alert
- [ ] Join incident channel (#incident-[number])
- [ ] Assess initial impact and scope
- [ ] Assign severity level
- [ ] Notify incident commander if SEV 1-2
### Investigation (15-60 minutes)
- [ ] Establish timeline of events
- [ ] Review recent changes (deployments, configs)
- [ ] Check dashboards and logs
- [ ] Identify affected components
- [ ] Communicate status update every 15 minutes
### Mitigation
- [ ] Identify mitigation options
- [ ] Evaluate risks of each option
- [ ] Implement chosen mitigation
- [ ] Verify mitigation effectiveness
- [ ] Update status page
### Resolution
- [ ] Confirm service restored
- [ ] Monitor for recurrence (30 min minimum)
- [ ] Send final status update
- [ ] Schedule postmortem
- [ ] Update incident ticket with timeline
Communication Templates
Initial Incident Communication:
**Incident Declared - [Service Name]**
Severity: SEV-[1/2/3]
Impact: [Brief description of user impact]
Start Time: [HH:MM UTC]
Status: Investigating
Current Actions:
- [What team is doing]
Next Update: [Time] or sooner if status changes
Incident Channel: #incident-[number]
Status Page: https://status.example.com
Status Update:
**Update - [Service Name] Incident**
Time: [HH:MM UTC]
Status: [Investigating/Identified/Monitoring/Resolved]
Summary:
[What happened since last update]
Current Actions:
[What team is doing now]
ETA to Resolution: [Time estimate if known]
Next Update: [Time]
Postmortem Templates
Key Concepts
Postmortem Principles:
- Blameless culture
- Focus on systemic improvements
- Share learnings broadly
- Track action items to completion
Postmortem Template
# Postmortem: [Incident Title]
**Date:** [YYYY-MM-DD]
**Duration:** [Start time] - [End time] ([total duration])
**Severity:** SEV-[1/2/3/4]
**Authors:** [Names]
**Status:** [Draft/In Review/Final]
## Executive Summary
[2-3 sentences describing what happened, impact, and resolution]
## Impact
- **User Impact:** [Percentage of users affected, specific user journeys impacted]
- **Revenue Impact:** [Estimated financial impact if applicable]
- **Duration:** [Total time users experienced impact]
- **Support Tickets:** [Number of tickets received]
## Timeline
All times in UTC.
| Time | Event |
|------|-------|
| 09:15 | Deployment of auth-service v2.3.1 completed |
| 09:23 | First alerts fire for elevated error rates |
| 09:25 | On-call engineer acknowledges alert |
| 09:30 | Incident declared as SEV-2 |
| 09:35 | Identified recent deployment as potential cause |
| 09:40 | Initiated rollback to v2.3.0 |
| 09:45 | Rollback completed |
| 09:50 | Error rates returning to normal |
| 10:20 | Incident resolved after 30 min monitoring |
## Root Cause Analysis
### What Happened
[Detailed technical explanation of the failure chain]
### Contributing Factors
1. **Factor 1:** [Description]
- Why it happened: [Explanation]
2. **Factor 2:** [Description]
- Why it happened: [Explanation]
### Five Whys Analysis
1. Why did users see errors?
- Because the auth service was returning 500 errors
2. Why was the auth service returning 500s?
- Because it couldn't connect to the database
3. Why couldn't it connect to the database?
- Because the connection pool was exhausted
4. Why was the connection pool exhausted?
- Because the new code had a connection leak
5. Why wasn't this caught in testing?
- Because load tests don't run long enough to detect slow leaks
## Detection
- **How was it detected:** [Automated alert / Customer report / Internal user]
- **Time to detect:** [Time from incident start to detection]
- **Detection gaps:** [What monitoring was missing or insufficient]
## Response
### What Went Well
- Alert fired within 2 minutes of impact starting
- On-call responded quickly
- Rollback procedure worked smoothly
- Communication was clear and timely
### What Could Be Improved
- Took 15 minutes to identify root cause
- Runbook was outdated for this scenario
- No automated rollback capability
## Action Items
| Priority | Action | Owner | Due Date | Status |
|----------|--------|-------|----------|--------|
| P1 | Add connection pool monitoring | @engineer1 | 2024-01-20 | In Progress |
| P1 | Implement automated rollback on error spike | @engineer2 | 2024-01-25 | Not Started |
| P2 | Extend load test duration to 4 hours | @engineer3 | 2024-01-22 | Not Started |
| P2 | Update auth-service runbook | @engineer1 | 2024-01-18 | Complete |
| P3 | Add connection leak detection to code review checklist | @tech-lead | 2024-01-30 | Not Started |
## Lessons Learned
1. **Lesson:** Connection pool metrics are critical for database-dependent services
- **Application:** Add pool utilisation alerts for all services
2. **Lesson:** Load tests need to simulate production duration
- **Application:** Create "soak test" category that runs for extended periods
3. **Lesson:** Automated rollback reduces MTTR significantly
- **Application:** Prioritise feature flag integration for critical services
## Supporting Information
- **Incident Ticket:** INCIDENT-1234
- **Related PRs:** #567, #568
- **Dashboards:** [Link to relevant dashboards]
- **Logs:** [Link to log queries]
- **Slack Channel:** #incident-20240115-auth-outage
Action Item Tracking
# Example action item tracking in YAML
postmortem_actions:
- id: PM-2024-001-01
title: "Add connection pool monitoring"
priority: P1
owner: "@engineer1"
due_date: "2024-01-20"
status: "in_progress"
verification: "Dashboard shows pool utilisation for all services"
- id: PM-2024-001-02
title: "Implement automated rollback"
priority: P1
owner: "@engineer2"
due_date: "2024-01-25"
status: "not_started"
dependencies:
- "Feature flag infrastructure"
verification: "Rollback triggers automatically when error rate > 5%"
Quick Reference
Observability Metrics Formulas
| Metric | Formula | Good Target |
|---|---|---|
| Availability | (Total - Errors) / Total | > 99.9% |
| Error Rate | Errors / Total | < 0.1% |
| P50 Latency | 50th percentile response time | < 100ms |
| P99 Latency | 99th percentile response time | < 500ms |
| Throughput | Requests / Time | Varies |
| MTTR | Sum(resolution times) / Incidents | < 1 hour |
| MTTA | Sum(acknowledgement times) / Alerts | < 5 minutes |
Common PromQL Queries
# Request rate
sum(rate(http_requests_total[5m])) by (service)
# Error rate percentage
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# P99 latency
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
# Availability over 30 days
avg_over_time(up{job="api"}[30d]) * 100
# Error budget burn rate
sum(rate(http_requests_total{status=~"5.."}[1h]))
/ sum(rate(http_requests_total[1h]))
/ (1 - 0.999) # Divide by error budget (1 - SLO)
Incident Severity Quick Guide
| Severity | User Impact | Response | Example |
|---|---|---|---|
| SEV 1 | Complete outage | Page all + exec | Site down |
| SEV 2 | Major degradation | Page team | Checkout broken |
| SEV 3 | Minor degradation | Working hours | Slow search |
| SEV 4 | Minimal | Next business day | Typo in UI |
Log Query Examples (Loki/LogQL)
# Errors in last hour
{service="api"} |= "error" | json | level="ERROR"
# Slow requests (> 1s)
{service="api"} | json | duration > 1000
# Requests by user
{service="api"} | json | user_id="user-123"
# Error rate by service
sum(rate({level="ERROR"}[5m])) by (service)
Common Issues and Solutions
Issue: Alert Fatigue
Symptoms:
- On-call ignores alerts
- High false positive rate
- Multiple alerts for same issue
Solutions:
- Implement alert inhibition rules
- Add
forduration to prevent flapping - Alert on symptoms, not causes
- Review and prune alerts quarterly
- Track alert actionability metrics
Issue: Missing Context in Alerts
Symptoms:
- Alerts don't provide enough information
- Engineers spend time gathering context
- Runbooks not linked
Solutions:
annotations:
summary: "Clear summary with current value: {{ $value }}"
description: |
**Impact:** Users cannot complete checkout
**Dashboard:** https://grafana.example.com/d/checkout
**Runbook:** https://wiki.example.com/runbooks/checkout-errors
**Recent Changes:** https://github.com/org/repo/commits
Issue: Logs Not Correlated with Traces
Symptoms:
- Cannot find logs for specific requests
- Debugging distributed issues is difficult
- Context lost between services
Solutions:
- Inject trace_id and span_id into all logs
- Use structured logging consistently
- Configure log aggregator to index trace fields
- Create dashboard links between trace and log views
Issue: SLO Not Reflecting User Experience
Symptoms:
- SLOs green but users complaining
- Metrics don't capture real issues
- Business stakeholders distrustful of data
Solutions:
- Define SLIs based on user journeys
- Use synthetic monitoring for critical paths
- Include client-side metrics
- Validate SLOs with real user feedback
- Regular SLO review meetings with stakeholders
Issue: Postmortems Not Leading to Improvements
Symptoms:
- Same incidents recurring
- Action items not completed
- Learnings not shared
Solutions:
- Track action items in issue tracker
- Review action item completion weekly
- Link postmortems to incident tickets
- Share postmortems in team meetings
- Create follow-up calendar reminders
Issue: On-Call Burnout
Symptoms:
- High interrupt rate
- Frequent night pages
- Engineer turnover
Solutions:
- Implement follow-the-sun rotations
- Set target for < 2 after-hours pages/week
- Invest in automation and self-healing
- Allow on-call to focus solely on incidents
- Compensate with time off or extra pay
Related Topics
The following topics would complement this Observability Patterns cheatsheet:
-
Prometheus & Grafana - Deep dive into metrics collection, PromQL queries, and dashboard design for comprehensive monitoring setup
-
Distributed Tracing with OpenTelemetry - Detailed coverage of instrumentation, context propagation, and trace analysis across microservices
-
Chaos Engineering - Proactive resilience testing patterns including fault injection, game days, and steady-state hypothesis
-
Site Reliability Engineering (SRE) Practices - Broader SRE principles including toil reduction, capacity planning, and production readiness
-
Log Management with ELK/Loki - Log aggregation architecture, query languages, retention policies, and cost optimisation
-
Kubernetes Monitoring - Container and orchestration-specific observability including resource metrics, cluster health, and workload monitoring