SLOs, SLIs, and Error Budgets
Service Level Objectives, Indicators, and Error Budget management for reliable systems.
SLOs, SLIs, and Error Budgets
Service Level Objectives, Indicators, and Error Budget management for reliable systems.
Overview
Service Level Indicators (SLIs) are quantitative measurements of service behaviour from a user's perspective. Service Level Objectives (SLOs) define target values or ranges for SLIs over a time window. Error Budgets quantify the acceptable amount of unreliability derived from SLOs, balancing reliability with innovation velocity.
This approach enables data-driven decisions about deployment freezes, incident response prioritisation, and feature development velocity based on actual user experience rather than arbitrary metrics.
graph TB
A[SLA - Service Level Agreement] --> B[SLO - Service Level Objective]
B --> C[SLI - Service Level Indicator]
A1[Legal Contract] --> A
B1[Internal Target] --> B
C1[Actual Measurement] --> C
B --> D[Error Budget]
D --> E[100% - SLO Target]
style A fill:#ff6b6b
style B fill:#4ecdc4
style C fill:#45b7d1
style D fill:#f9ca24
SLI Fundamentals
Defining User-Centric SLIs
Focus on what users actually experience, not internal system metrics.
Good SLIs:
- Request success rate (availability)
- Request latency at percentiles (speed)
- Throughput (capacity)
- Data freshness (consistency)
Poor SLIs:
- CPU utilisation
- Memory usage
- Queue depth
- Internal error rates users don't see
SLI Specification Structure
Every SLI should have:
- Indicator: What you're measuring
- Measurement: How you measure it
- Valid events: What counts as a measurement
- Good events: What counts as success
# Example SLI Specification
availability_sli:
indicator: "Request Success Rate"
measurement: "Ratio of successful requests to total requests"
valid_events: "All HTTP requests to /api/* endpoints"
good_events: "HTTP responses with status 200-299, 304, or 401-403"
time_window: "28 days"
# Note: 401-403 (and 4xx generally) are auth/client errors, not server faults,
# so they're treated as "good" for an availability SLI - the service responded
# correctly. Decide per service whether to count or exclude them; don't copy blindly.
Common SLI Types
Availability SLI
Measures the proportion of successful requests.
# Prometheus query for availability SLI
sum(rate(http_requests_total{status=~"2..|304|401|402|403"}[5m]))
/
sum(rate(http_requests_total[5m]))
Calculation:
Availability = Good Events / Valid Events
= Successful Requests / Total Requests
Latency SLI
Measures request duration at specific percentiles.
# 95th percentile latency over 5 minutes
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# Proportion of requests faster than 300ms
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))
Common thresholds:
- Fast: < 100ms
- Good: < 300ms
- Acceptable: < 1000ms
- Slow: > 1000ms
Throughput SLI
Measures the proportion of time the system can handle expected load.
# Proportion of time processing > minimum threshold requests/sec
(
sum(rate(http_requests_total[1m])) > 100
) / 60
SLI Implementation
# Python example: Calculate availability SLI
from prometheus_api_client import PrometheusConnect
def calculate_availability_sli(prometheus_url, window_hours=1):
"""Calculate availability SLI from Prometheus metrics."""
prom = PrometheusConnect(url=prometheus_url)
query = f'''
sum(rate(http_requests_total{{status=~"2..|304"}}[{window_hours}h]))
/
sum(rate(http_requests_total[{window_hours}h]))
'''
result = prom.custom_query(query=query)
if result:
availability = float(result[0]['value'][1])
return availability * 100 # Return as percentage
return None
# Example usage
availability = calculate_availability_sli('http://prometheus:9090', window_hours=24)
print(f"24-hour availability: {availability:.3f}%")
SLO Target Setting
Choosing SLO Targets
flowchart TD
A[Start with User Expectations] --> B{Historical Performance}
B -->|Good| C[Set SLO at 99th percentile of current performance]
B -->|Poor| D[Set aspirational SLO + phased improvement plan]
C --> E{Business Requirements}
D --> E
E -->|Critical Service| F[Higher SLO: 99.9% - 99.99%]
E -->|Standard Service| G[Standard SLO: 99% - 99.5%]
E -->|Batch/Internal| H[Lower SLO: 95% - 99%]
F --> I[Calculate Error Budget]
G --> I
H --> I
I --> J[Validate with Stakeholders]
J --> K{Achievable?}
K -->|No| L[Adjust SLO or Invest in Reliability]
K -->|Yes| M[Document and Implement]
L --> E
SLO Target Examples
| Service Type | Availability SLO | Latency SLO (p95) | Error Budget/month |
|---|---|---|---|
| Critical user-facing API | 99.95% | 200ms | 21.6 minutes |
| Standard web service | 99.9% | 500ms | 43.2 minutes |
| Internal service | 99.5% | 1000ms | 3.6 hours |
| Batch processing | 99% | N/A | 7.2 hours |
| Background jobs | 95% | N/A | 36 hours |
Calculation Windows
Choose windows based on user impact and operational practicality:
Rolling Windows:
# 30-day rolling window
slo:
target: 99.9
window: 30d
# Pros: Smooth, no reset cliff
# Cons: Slower to recover from incidents
Calendar Windows:
# Monthly calendar window
slo:
target: 99.5
window: calendar_month
# Pros: Aligns with business cycles, fresh start each period
# Cons: End-of-period gaming, sudden resets
Multiple Windows:
# Recommended: Use both short and long windows
slos:
- window: 28d
target: 99.9
- window: 7d
target: 99.5
# Short window: Catch recent trends
# Long window: Ensure sustained reliability
Error Budgets
Error Budget Calculation
flowchart LR
A[SLO Target: 99.9%] --> B[Acceptable Downtime: 0.1%]
B --> C[Error Budget = 100% - 99.9%]
C --> D[Error Budget = 0.1%]
D --> E[Convert to Time]
E --> F[Per Day: 86.4 seconds]
E --> G[Per Month: 43.2 minutes]
E --> H[Per Year: 8.76 hours]
Error Budget Examples
Example 1: Availability-based Error Budget
SLO: 99.95% availability over 30 days
Total requests in 30 days: 100,000,000
Error budget: 100% - 99.95% = 0.05%
Allowed failed requests: 100,000,000 × 0.05% = 50,000
Example 2: Latency-based Error Budget
SLO: 95% of requests < 300ms over 7 days
Total requests in 7 days: 10,000,000
Error budget: 100% - 95% = 5%
Allowed slow requests: 10,000,000 × 5% = 500,000
Example 3: Combined Error Budget
Service with multiple SLOs:
- Availability: 99.9% (0.1% error budget)
- Latency p95: < 200ms (5% can be slower)
If 1M requests/day:
- Can fail: 1,000 requests/day
- Can be slow: 50,000 requests/day
Error Budget Tracking
# Prometheus query for error budget consumption
# Calculate error budget remaining (availability-based)
1 - (
(1 - sum(rate(http_requests_total{status=~"2..|304"}[30d]))
/ sum(rate(http_requests_total[30d])))
/
(1 - 0.999) # SLO target: 99.9%
)
# Result:
# 1.0 = 100% budget remaining
# 0.5 = 50% budget remaining
# 0.0 = 0% budget exhausted
# Python: Calculate error budget status
def calculate_error_budget_status(current_sli, slo_target, window_days=30):
"""
Calculate error budget consumption.
Args:
current_sli: Current SLI value (e.g., 0.9995 for 99.95%)
slo_target: SLO target (e.g., 0.999 for 99.9%)
window_days: Rolling window in days
Returns:
dict with error budget metrics
"""
error_budget = 1 - slo_target
actual_errors = 1 - current_sli
consumed = actual_errors / error_budget if error_budget > 0 else 0
remaining = max(0, 1 - consumed)
# Time calculations
window_minutes = window_days * 24 * 60
budget_minutes = window_minutes * error_budget
consumed_minutes = window_minutes * actual_errors
remaining_minutes = budget_minutes - consumed_minutes
return {
'budget_remaining_pct': remaining * 100,
'budget_consumed_pct': consumed * 100,
'remaining_minutes': max(0, remaining_minutes),
'consumed_minutes': consumed_minutes,
'total_budget_minutes': budget_minutes,
'status': 'healthy' if consumed < 0.8 else 'warning' if consumed < 1.0 else 'exhausted'
}
# Example usage
status = calculate_error_budget_status(
current_sli=0.9992, # 99.92% actual
slo_target=0.999, # 99.9% target
window_days=30
)
print(f"Error budget: {status['budget_remaining_pct']:.1f}% remaining ({status['status']})")
# Output: Error budget: 20.0% remaining (healthy)
Error Budget Policies
Policy Framework
# Example error budget policy
error_budget_policy:
service: "payment-api"
slo_target: 99.9%
measurement_window: 28d
actions:
- threshold: 100%
status: "exhausted"
actions:
- "Incident declared automatically"
- "Feature releases blocked"
- "All hands focus on reliability"
- "Executive escalation"
- threshold: 80%
status: "warning"
actions:
- "Increase monitoring frequency"
- "Review pending deployments for risk"
- "Defer non-critical releases"
- "Daily reliability review"
- threshold: 50%
status: "caution"
actions:
- "Heightened awareness"
- "Additional testing for releases"
- threshold: 0%
status: "healthy"
actions:
- "Normal operations"
- "Consider investing in new features"
Decision Framework
flowchart TD
A[Error Budget Check] --> B{Budget Status?}
B -->|> 50% Remaining| C[Healthy]
B -->|20-50% Remaining| D[Caution]
B -->|0-20% Remaining| E[Warning]
B -->|Exhausted| F[Critical]
C --> C1[Normal deployment velocity]
C --> C2[Consider feature work]
D --> D1[Extra testing required]
D --> D2[Review deployment risk]
E --> E1[Reduce deployment frequency]
E --> E2[Focus on reliability]
E --> E3[Daily reviews]
F --> F1[Block feature releases]
F --> F2[All hands on reliability]
F --> F3[Incident declared]
style C fill:#90EE90
style D fill:#FFD700
style E fill:#FFA500
style F fill:#FF6347
Burn Rate Alerts
Burn Rate Concepts
Burn rate measures how quickly you're consuming error budget. A burn rate of 1 means you're consuming budget at exactly the rate that would exhaust it by the end of the window.
Burn Rate = (Current Error Rate) / (Maximum Allowed Error Rate)
Example:
SLO: 99.9% (0.1% error budget)
Current error rate: 1% (10× the budget)
Burn rate: 1% / 0.1% = 10
At this rate, error budget exhausted in: 30 days / 10 = 3 days
Multi-Window, Multi-Burn-Rate Alerting
flowchart TD
A[Error Budget Monitoring] --> B[Fast Burn Alert]
A --> C[Moderate Burn Alert]
A --> D[Slow Burn Alert]
B --> B1[1-hour window]
B --> B2[Burn rate > 14.4]
B --> B3[Budget depleted in < 2 days]
B --> B4[Page immediately]
C --> C1[6-hour window]
C --> C2[Burn rate > 6]
C --> C3[Budget depleted in < 5 days]
C --> C4[Alert during business hours]
D --> D1[3-day window]
D --> D2[Burn rate > 1]
D --> D3[Budget depleted by end of month]
D --> D4[Ticket for investigation]
style B4 fill:#ff6b6b
style C4 fill:#ffa500
style D4 fill:#ffd700
Alert Configuration Examples
Fast Burn (Page):
# High severity: Budget exhausted in < 2 days
alert: SLOBurnRateFast
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
) > (14.4 * (1 - 0.999))
and
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
) > (14.4 * (1 - 0.999))
for: 2m
severity: page
labels:
burn_rate: "fast"
depletion_time: "2d"
annotations:
summary: "Fast error budget burn: exhaustion in < 2 days"
Moderate Burn (Alert):
# Medium severity: Budget exhausted in < 5 days
alert: SLOBurnRateModerate
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[6h]))
/
sum(rate(http_requests_total[6h]))
) > (6 * (1 - 0.999))
and
(
sum(rate(http_requests_total{status=~"5.."}[30m]))
/
sum(rate(http_requests_total[30m]))
) > (6 * (1 - 0.999))
for: 15m
severity: alert
labels:
burn_rate: "moderate"
depletion_time: "5d"
Slow Burn (Ticket):
# Low severity: Budget exhausted by end of window
alert: SLOBurnRateSlow
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[3d]))
/
sum(rate(http_requests_total[3d]))
) > (1 * (1 - 0.999))
and
(
sum(rate(http_requests_total{status=~"5.."}[6h]))
/
sum(rate(http_requests_total[6h]))
) > (1 * (1 - 0.999))
for: 1h
severity: ticket
labels:
burn_rate: "slow"
depletion_time: "30d"
Burn Rate Thresholds
For a 30-day SLO window:
| Alert Type | Window | Burn Rate | Time to Exhaustion | Action |
|---|---|---|---|---|
| Critical | 1 hour | 14.4× | < 2 days | Page on-call |
| High | 6 hours | 6× | < 5 days | Alert team |
| Medium | 1 day | 3× | < 10 days | Create ticket |
| Low | 3 days | 1× | 30 days | Monitor |
Calculation for burn rate thresholds:
Burn Rate = (Window Duration) / (Acceptable Depletion Time)
For 2-day depletion on 30-day window:
Burn Rate = 30 / 2 = 15 (rounded to 14.4 for 5% budget consumption)
Instrumentation
Availability Instrumentation
Application-level tracking:
# Python Flask example with Prometheus
from flask import Flask, request
from prometheus_client import Counter, Histogram, generate_latest
import time
app = Flask(__name__)
# Metrics
request_count = Counter(
'http_requests_total',
'Total HTTP requests',
['method', 'endpoint', 'status']
)
request_latency = Histogram(
'http_request_duration_seconds',
'HTTP request latency',
['method', 'endpoint']
)
@app.before_request
def before_request():
request.start_time = time.time()
@app.after_request
def after_request(response):
# Record request count
request_count.labels(
method=request.method,
endpoint=request.endpoint or 'unknown',
status=response.status_code
).inc()
# Record latency
if hasattr(request, 'start_time'):
duration = time.time() - request.start_time
request_latency.labels(
method=request.method,
endpoint=request.endpoint or 'unknown'
).observe(duration)
return response
@app.route('/api/users')
def get_users():
# Your application logic
return {'users': []}
@app.route('/metrics')
def metrics():
return generate_latest()
Nginx instrumentation:
# nginx.conf with request tracking
http {
log_format sli_format '$remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'$request_time $upstream_response_time '
'"$http_user_agent"';
access_log /var/log/nginx/access.log sli_format;
# Export metrics to Prometheus
server {
location /metrics {
stub_status on;
access_log off;
}
}
}
Latency Instrumentation
Using Prometheus histograms:
# Python: Latency tracking with custom buckets
from prometheus_client import Histogram
# Define buckets relevant to your SLOs
http_latency = Histogram(
'http_request_duration_seconds',
'HTTP request latency in seconds',
['method', 'endpoint'],
buckets=[0.1, 0.3, 0.5, 1.0, 2.0, 5.0, 10.0] # SLO-aligned buckets
)
# Usage
with http_latency.labels(method='GET', endpoint='/api/users').time():
# Your application logic
process_request()
Node.js Express example:
// Express middleware for SLI tracking
const promClient = require('prom-client');
const httpRequestDuration = new promClient.Histogram({
name: 'http_request_duration_seconds',
help: 'Duration of HTTP requests in seconds',
labelNames: ['method', 'route', 'status_code'],
buckets: [0.1, 0.3, 0.5, 1.0, 2.0, 5.0]
});
const httpRequestTotal = new promClient.Counter({
name: 'http_requests_total',
help: 'Total number of HTTP requests',
labelNames: ['method', 'route', 'status_code']
});
app.use((req, res, next) => {
const start = Date.now();
res.on('finish', () => {
const duration = (Date.now() - start) / 1000;
httpRequestDuration.labels(
req.method,
req.route?.path || 'unknown',
res.statusCode
).observe(duration);
httpRequestTotal.labels(
req.method,
req.route?.path || 'unknown',
res.statusCode
).inc();
});
next();
});
Prometheus Recording Rules
Pre-aggregate SLI calculations for efficient querying:
# prometheus_rules.yml
groups:
- name: sli_rules
interval: 30s
rules:
# Availability SLI
- record: sli:availability:ratio_rate5m
expr: |
sum(rate(http_requests_total{status=~"2..|304|401|402|403"}[5m]))
/
sum(rate(http_requests_total[5m]))
# Latency SLI (proportion fast)
- record: sli:latency:ratio_rate5m
expr: |
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))
# Error budget consumption (30-day)
- record: sli:error_budget:consumed_ratio_30d
expr: |
1 - (
(1 - sli:availability:ratio_rate5m)
/
(1 - 0.999) # SLO target
)
Reporting and Review
SLO Dashboard Components
Key metrics to display:
- Current SLI value (real-time)
- SLO target (static reference)
- Error budget remaining (percentage and time)
- Burn rate (current rate)
- Trend graph (30-day history)
- Time to exhaustion (at current burn rate)
Grafana dashboard JSON snippet:
{
"panels": [
{
"title": "Availability SLI",
"targets": [
{
"expr": "sli:availability:ratio_rate5m * 100"
}
],
"thresholds": [
{"value": 99.9, "color": "red"},
{"value": 99.95, "color": "yellow"},
{"value": 100, "color": "green"}
]
},
{
"title": "Error Budget Remaining",
"targets": [
{
"expr": "sli:error_budget:consumed_ratio_30d * 100"
}
],
"gauge": {
"maxValue": 100,
"minValue": 0,
"thresholds": [
{"value": 0, "color": "red"},
{"value": 20, "color": "orange"},
{"value": 50, "color": "yellow"},
{"value": 80, "color": "green"}
]
}
}
]
}
Review Cadences
# Recommended review schedule
slo_review_schedule:
daily:
- audience: "Engineering team"
- review: "Error budget status"
- action: "Adjust deployment velocity if needed"
- duration: "5 minutes in standup"
weekly:
- audience: "Service owners + SRE"
- review: "SLO compliance, trends, incidents"
- action: "Identify reliability improvements"
- duration: "30 minutes"
monthly:
- audience: "Engineering + Product + Leadership"
- review: "SLO performance, error budget usage, policy effectiveness"
- action: "Adjust SLOs if needed, prioritise reliability work"
- duration: "1 hour"
quarterly:
- audience: "All stakeholders"
- review: "SLO strategy, targets, new services"
- action: "Set reliability goals for next quarter"
- duration: "2 hours"
Review Meeting Template
# Weekly SLO Review - [Date]
## Services Reviewed
- Service A (payment-api)
- Service B (user-service)
- Service C (notification-service)
## SLO Compliance
| Service | SLO Target | Current SLI | Status | Error Budget |
|---------|-----------|-------------|--------|--------------|
| payment-api | 99.9% | 99.92% | ✅ Healthy | 20% remaining |
| user-service | 99.95% | 99.89% | ⚠️ Warning | -120% (exhausted) |
| notification-service | 99.5% | 99.87% | ✅ Healthy | 74% remaining |
## Incidents Impact
- INC-1234: Database failover (user-service) - consumed 80% of monthly budget
- INC-1235: API timeout spike (payment-api) - consumed 15% of monthly budget
## Actions
1. user-service: Release freeze until error budget recovers
2. user-service: Root cause analysis scheduled for tomorrow
3. payment-api: Continue monitoring, normal operations
## Reliability Improvements
- Implement circuit breaker for user-service database calls
- Add retry logic with exponential backoff for payment-api
Complete Examples
Example 1: E-commerce Checkout SLO
service: checkout-api
description: "Payment processing and order completion"
slis:
- name: availability
description: "Proportion of successful checkout requests"
measurement: |
sum(rate(http_requests_total{service="checkout",status=~"2.."}[5m]))
/
sum(rate(http_requests_total{service="checkout"}[5m]))
valid_events: "All POST /api/checkout requests"
good_events: "HTTP 200-299 responses"
- name: latency
description: "Proportion of fast checkout requests"
measurement: |
sum(rate(http_request_duration_seconds_bucket{service="checkout",le="2.0"}[5m]))
/
sum(rate(http_request_duration_seconds_count{service="checkout"}[5m]))
valid_events: "All POST /api/checkout requests"
good_events: "Requests completed in < 2 seconds"
slos:
- sli: availability
target: 99.95%
window: 30d
- sli: latency
target: 99.0%
window: 30d
error_budget_policy:
- remaining: 100%
actions: ["Deployment freeze", "Incident declared", "All hands"]
- remaining: 80%
actions: ["Reduce deployment frequency", "Extra testing"]
- remaining: 50%
actions: ["Heightened monitoring"]
alerts:
- name: CheckoutAvailabilityBurnFast
expr: |
(
1 - sum(rate(http_requests_total{service="checkout",status=~"2.."}[1h]))
/ sum(rate(http_requests_total{service="checkout"}[1h]))
) > (14.4 * 0.0005)
severity: page
- name: CheckoutLatencyBurnModerate
expr: |
(
1 - sum(rate(http_request_duration_seconds_bucket{service="checkout",le="2.0"}[6h]))
/ sum(rate(http_request_duration_seconds_count{service="checkout"}[6h]))
) > (6 * 0.01)
severity: alert
Example 2: Data Pipeline SLO
service: data-ingestion-pipeline
description: "Real-time event processing pipeline"
slis:
- name: freshness
description: "Proportion of events processed within SLA"
measurement: |
sum(rate(events_processed{lag_seconds<300}[5m]))
/
sum(rate(events_total[5m]))
valid_events: "All events received"
good_events: "Events processed within 5 minutes"
- name: correctness
description: "Proportion of events processed without errors"
measurement: |
sum(rate(events_processed{status="success"}[5m]))
/
sum(rate(events_processed[5m]))
valid_events: "All processing attempts"
good_events: "Successfully processed events"
slos:
- sli: freshness
target: 99.0%
window: 7d
- sli: correctness
target: 99.9%
window: 7d
error_budget_policy:
- remaining: 100%
actions: ["Pause data source integration", "Investigate immediately"]
- remaining: 50%
actions: ["Monitor closely", "Prepare rollback plan"]
Quick Reference
SLO Target Selection
| Service Criticality | Availability | Latency (p95) | Window |
|---|---|---|---|
| Critical user-facing | 99.95% - 99.99% | 100-300ms | 28-30d |
| Standard user-facing | 99.5% - 99.9% | 300-1000ms | 28-30d |
| Internal service | 99% - 99.5% | 1000-3000ms | 7-28d |
| Batch/Background | 95% - 99% | N/A | 7-28d |
Error Budget Time Allowances
| SLO Target | Downtime/30 days | Downtime/year |
|---|---|---|
| 99.99% | 4.32 minutes | 52.6 minutes |
| 99.95% | 21.6 minutes | 4.38 hours |
| 99.9% | 43.2 minutes | 8.76 hours |
| 99.5% | 3.6 hours | 1.83 days |
| 99% | 7.2 hours | 3.65 days |
| 95% | 36 hours | 18.25 days |
Common Prometheus Queries
# Availability SLI (5-minute window)
sum(rate(http_requests_total{status=~"2..|304"}[5m]))
/
sum(rate(http_requests_total[5m]))
# Latency SLI (proportion fast, 5-minute window)
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))
# Error budget remaining (30-day window, 99.9% target)
1 - (
(1 - sli:availability:ratio_rate5m)
/
(1 - 0.999)
)
# Burn rate (1-hour window)
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
)
/
(1 - 0.999)
# Time to error budget exhaustion (hours)
(
sli:error_budget:consumed_ratio_30d * (30 * 24)
)
/
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
)
Burn Rate Alert Thresholds (30-day window)
| Window | Burn Rate | Exhaustion Time | Severity |
|---|---|---|---|
| 1 hour | 14.4× | 2 days | Page |
| 6 hours | 6× | 5 days | Alert |
| 1 day | 3× | 10 days | Warn |
| 3 days | 1× | 30 days | Info |
Error Budget Policy Actions
| Budget Remaining | Status | Actions |
|---|---|---|
| > 50% | Healthy | Normal velocity, consider feature work |
| 20-50% | Caution | Extra testing, review risky changes |
| 0-20% | Warning | Reduce deployments, daily reviews, focus on reliability |
| Exhausted | Critical | Freeze features, all hands on reliability, incident declared |
Common Issues and Solutions
Issue: SLI doesn't reflect user experience
Symptoms:
- Users report problems but SLI shows green
- SLI is 100% but users experience errors
Solutions:
# Add client-side measurements
- Implement Real User Monitoring (RUM)
- Track synthetic probes from user locations
- Include mobile app metrics
- Monitor from outside your network
# Example: Add synthetic monitoring
apiVersion: v1
kind: ConfigMap
metadata:
name: blackbox-exporter
data:
config.yml: |
modules:
http_2xx:
prober: http
timeout: 5s
http:
valid_status_codes: [200, 201, 204]
fail_if_not_ssl: true
preferred_ip_protocol: "ip4"
Issue: Error budget exhausted too quickly
Symptoms:
- Frequently hitting error budget limits
- Constant deployment freezes
- Team velocity severely impacted
Solutions:
-
Re-evaluate SLO target - May be too aggressive
Current: 99.99% (4.32 min/month) Proposed: 99.95% (21.6 min/month) - 5× more budget -
Exclude expected failures from SLI
# Don't count 4xx client errors (except 429) in availability sum(rate(http_requests_total{status=~"2..|304|4.."}[5m])) -
Implement gradual rollouts to reduce blast radius
deployment_strategy: - canary: 5% duration: 30m - canary: 25% duration: 1h - canary: 100%
Issue: Alert fatigue from burn rate alerts
Symptoms:
- Too many burn rate alerts
- Alerts during normal operations
- Team ignoring alerts
Solutions:
# Adjust burn rate thresholds
# Before: Alert at 6× burn rate (too sensitive)
# After: Alert at 10× burn rate
# Add multi-window confirmation
- expr: |
(error_rate[1h] > threshold) # Short window
and
(error_rate[6h] > threshold) # Long window confirmation
for: 15m # Must persist for 15 minutes
# Reduce noise during deployments
- expr: |
slo_burn_rate > 10
and
absent(deployment_in_progress{service="api"})
Issue: Cannot meet SLO with dependencies
Symptoms:
- Your SLO: 99.9%
- Dependency SLO: 99.5%
- Can't achieve target
Solutions:
# 1. Adjust your SLO based on dependency chain
# If you depend on 3 services at 99.5% each:
# Maximum achievable: 99.5% × 99.5% × 99.5% = 98.5%
# Set your SLO realistically: 98.0%
# 2. Implement resilience patterns
patterns:
- circuit_breaker:
failure_threshold: 5
timeout: 30s
recovery_time: 60s
- retry_with_backoff:
max_attempts: 3
initial_delay: 100ms
multiplier: 2
- fallback:
cache: stale_data
timeout: 5s
- timeout:
connection: 1s
request: 5s
# 3. Cache aggressively
cache_strategy:
ttl: 300s
stale_while_revalidate: 600s
serve_stale_on_error: true
Issue: Latency SLI gaming
Symptoms:
- System terminates slow requests to meet SLO
- Users see more failures but latency SLI looks good
Solutions:
# Use both availability AND latency SLOs
slos:
- name: availability
target: 99.9% # Can't just drop requests
- name: latency
target: 99.0% # Must be fast AND successful
# Track all outcomes
sli_definition:
valid_events: "All requests initiated by users"
good_events: "Requests that completed successfully AND within latency target"
# Bad approach: Excludes timeouts
# good_events: "Successful requests < 300ms"
# Good approach: Includes all outcomes
# Timeout counts as both availability AND latency failure
Issue: Different user expectations by region
Symptoms:
- Global SLO doesn't reflect regional experience
- Some regions consistently poor
Solutions:
# Define SLOs per region
slos:
- region: us-east
availability: 99.95%
latency_p95: 100ms
- region: eu-west
availability: 99.95%
latency_p95: 150ms
- region: ap-southeast
availability: 99.9%
latency_p95: 300ms
# Prometheus query with regional labels
sum(rate(http_requests_total{status=~"2..",region="us-east"}[5m]))
/
sum(rate(http_requests_total{region="us-east"}[5m]))
Issue: Monthly SLO resets create perverse incentives
Symptoms:
- Team "saves" error budget for end of month
- Riskier deploys at month start
- Sudden focus on reliability at month end
Solutions:
# Use rolling windows instead of calendar windows
slo:
window: 28d # Rolling 28-day window
target: 99.9%
# Or use multiple windows
slos:
- window: 7d # Short-term quality signal
target: 99.5%
- window: 28d # Long-term trend
target: 99.9%
# Both must be met
Issue: Can't identify which component caused SLO violation
Symptoms:
- SLO alert fires but unclear which service/component is responsible
- Long investigation time
Solutions:
# Add detailed labels to track attribution
http_requests_total{
service="api",
component="database",
operation="query",
failure_mode="timeout"
}
# Create SLIs per critical component
- sli:availability:database
- sli:availability:cache
- sli:availability:external_api
# Use distributed tracing
# Tag requests with trace_id and analyse:
- Which service added the most latency?
- Which component had the error?
- What was the failure propagation path?