Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

SLOs, SLIs, and Error Budgets

Service Level Objectives, Indicators, and Error Budget management for reliable systems.

SLOs, SLIs, and Error Budgets

Service Level Objectives, Indicators, and Error Budget management for reliable systems.

Overview

Service Level Indicators (SLIs) are quantitative measurements of service behaviour from a user's perspective. Service Level Objectives (SLOs) define target values or ranges for SLIs over a time window. Error Budgets quantify the acceptable amount of unreliability derived from SLOs, balancing reliability with innovation velocity.

This approach enables data-driven decisions about deployment freezes, incident response prioritisation, and feature development velocity based on actual user experience rather than arbitrary metrics.

SLA - Service LevelAgreementSLO - Service LevelObjectiveSLI - Service LevelIndicatorLegal ContractInternal TargetActual MeasurementError Budget100% - SLO TargetSLA - Service LevelAgreementSLO - Service LevelObjectiveSLI - Service LevelIndicatorLegal ContractInternal TargetActual MeasurementError Budget100% - SLO Target

SLI Fundamentals

Defining User-Centric SLIs

Focus on what users actually experience, not internal system metrics.

Good SLIs:

  • Request success rate (availability)
  • Request latency at percentiles (speed)
  • Throughput (capacity)
  • Data freshness (consistency)

Poor SLIs:

  • CPU utilisation
  • Memory usage
  • Queue depth
  • Internal error rates users don't see

SLI Specification Structure

Every SLI should have:

  1. Indicator: What you're measuring
  2. Measurement: How you measure it
  3. Valid events: What counts as a measurement
  4. Good events: What counts as success
# Example SLI Specification
availability_sli:
  indicator: "Request Success Rate"
  measurement: "Ratio of successful requests to total requests"
  valid_events: "All HTTP requests to /api/* endpoints"
  good_events: "HTTP responses with status 200-299, 304, or 401-403"
  time_window: "28 days"
  # Note: 401-403 (and 4xx generally) are auth/client errors, not server faults,
  # so they're treated as "good" for an availability SLI - the service responded
  # correctly. Decide per service whether to count or exclude them; don't copy blindly.

Common SLI Types

Availability SLI

Measures the proportion of successful requests.

# Prometheus query for availability SLI
sum(rate(http_requests_total{status=~"2..|304|401|402|403"}[5m]))
/
sum(rate(http_requests_total[5m]))

Calculation:

Availability = Good Events / Valid Events
             = Successful Requests / Total Requests

Latency SLI

Measures request duration at specific percentiles.

# 95th percentile latency over 5 minutes
histogram_quantile(0.95,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)

# Proportion of requests faster than 300ms
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))

Common thresholds:

  • Fast: < 100ms
  • Good: < 300ms
  • Acceptable: < 1000ms
  • Slow: > 1000ms

Throughput SLI

Measures the proportion of time the system can handle expected load.

# Proportion of time processing > minimum threshold requests/sec
(
  sum(rate(http_requests_total[1m])) > 100
) / 60

SLI Implementation

# Python example: Calculate availability SLI
from prometheus_api_client import PrometheusConnect

def calculate_availability_sli(prometheus_url, window_hours=1):
    """Calculate availability SLI from Prometheus metrics."""
    prom = PrometheusConnect(url=prometheus_url)

    query = f'''
    sum(rate(http_requests_total{{status=~"2..|304"}}[{window_hours}h]))
    /
    sum(rate(http_requests_total[{window_hours}h]))
    '''

    result = prom.custom_query(query=query)

    if result:
        availability = float(result[0]['value'][1])
        return availability * 100  # Return as percentage
    return None

# Example usage
availability = calculate_availability_sli('http://prometheus:9090', window_hours=24)
print(f"24-hour availability: {availability:.3f}%")

SLO Target Setting

Choosing SLO Targets

GoodPoorCritical ServiceStandard ServiceBatch/InternalNoYesStart with UserExpectationsHistoricalPerformanceSet SLO at 99thpercentile ofcurrent performanceSet aspirational SLO+ phased improvementplanBusinessRequirementsHigher SLO: 99.9% -99.99%Standard SLO: 99% -99.5%Lower SLO: 95% - 99%Calculate ErrorBudgetValidate withStakeholdersAchievable?Adjust SLO or Investin ReliabilityDocument andImplementGoodPoorCritical ServiceStandard ServiceBatch/InternalNoYesStart with UserExpectationsHistoricalPerformanceSet SLO at 99thpercentile ofcurrent performanceSet aspirational SLO+ phased improvementplanBusinessRequirementsHigher SLO: 99.9% -99.99%Standard SLO: 99% -99.5%Lower SLO: 95% - 99%Calculate ErrorBudgetValidate withStakeholdersAchievable?Adjust SLO or Investin ReliabilityDocument andImplement

SLO Target Examples

Service Type Availability SLO Latency SLO (p95) Error Budget/month
Critical user-facing API 99.95% 200ms 21.6 minutes
Standard web service 99.9% 500ms 43.2 minutes
Internal service 99.5% 1000ms 3.6 hours
Batch processing 99% N/A 7.2 hours
Background jobs 95% N/A 36 hours

Calculation Windows

Choose windows based on user impact and operational practicality:

Rolling Windows:

# 30-day rolling window
slo:
  target: 99.9
  window: 30d

# Pros: Smooth, no reset cliff
# Cons: Slower to recover from incidents

Calendar Windows:

# Monthly calendar window
slo:
  target: 99.5
  window: calendar_month

# Pros: Aligns with business cycles, fresh start each period
# Cons: End-of-period gaming, sudden resets

Multiple Windows:

# Recommended: Use both short and long windows
slos:
  - window: 28d
    target: 99.9
  - window: 7d
    target: 99.5

# Short window: Catch recent trends
# Long window: Ensure sustained reliability

Error Budgets

Error Budget Calculation

SLO Target: 99.9%Acceptable Downtime:0.1%Error Budget = 100%- 99.9%Error Budget = 0.1%Convert to TimePer Day: 86.4secondsPer Month: 43.2minutesPer Year: 8.76 hoursSLO Target: 99.9%Acceptable Downtime:0.1%Error Budget = 100%- 99.9%Error Budget = 0.1%Convert to TimePer Day: 86.4secondsPer Month: 43.2minutesPer Year: 8.76 hours

Error Budget Examples

Example 1: Availability-based Error Budget

SLO: 99.95% availability over 30 days
Total requests in 30 days: 100,000,000
Error budget: 100% - 99.95% = 0.05%

Allowed failed requests: 100,000,000 × 0.05% = 50,000

Example 2: Latency-based Error Budget

SLO: 95% of requests < 300ms over 7 days
Total requests in 7 days: 10,000,000
Error budget: 100% - 95% = 5%

Allowed slow requests: 10,000,000 × 5% = 500,000

Example 3: Combined Error Budget

Service with multiple SLOs:
- Availability: 99.9% (0.1% error budget)
- Latency p95: < 200ms (5% can be slower)

If 1M requests/day:
- Can fail: 1,000 requests/day
- Can be slow: 50,000 requests/day

Error Budget Tracking

# Prometheus query for error budget consumption
# Calculate error budget remaining (availability-based)

1 - (
  (1 - sum(rate(http_requests_total{status=~"2..|304"}[30d]))
       / sum(rate(http_requests_total[30d])))
  /
  (1 - 0.999)  # SLO target: 99.9%
)

# Result:
# 1.0 = 100% budget remaining
# 0.5 = 50% budget remaining
# 0.0 = 0% budget exhausted
# Python: Calculate error budget status
def calculate_error_budget_status(current_sli, slo_target, window_days=30):
    """
    Calculate error budget consumption.

    Args:
        current_sli: Current SLI value (e.g., 0.9995 for 99.95%)
        slo_target: SLO target (e.g., 0.999 for 99.9%)
        window_days: Rolling window in days

    Returns:
        dict with error budget metrics
    """
    error_budget = 1 - slo_target
    actual_errors = 1 - current_sli

    consumed = actual_errors / error_budget if error_budget > 0 else 0
    remaining = max(0, 1 - consumed)

    # Time calculations
    window_minutes = window_days * 24 * 60
    budget_minutes = window_minutes * error_budget
    consumed_minutes = window_minutes * actual_errors
    remaining_minutes = budget_minutes - consumed_minutes

    return {
        'budget_remaining_pct': remaining * 100,
        'budget_consumed_pct': consumed * 100,
        'remaining_minutes': max(0, remaining_minutes),
        'consumed_minutes': consumed_minutes,
        'total_budget_minutes': budget_minutes,
        'status': 'healthy' if consumed < 0.8 else 'warning' if consumed < 1.0 else 'exhausted'
    }

# Example usage
status = calculate_error_budget_status(
    current_sli=0.9992,  # 99.92% actual
    slo_target=0.999,    # 99.9% target
    window_days=30
)
print(f"Error budget: {status['budget_remaining_pct']:.1f}% remaining ({status['status']})")
# Output: Error budget: 20.0% remaining (healthy)

Error Budget Policies

Policy Framework

# Example error budget policy
error_budget_policy:
  service: "payment-api"
  slo_target: 99.9%
  measurement_window: 28d

  actions:
    - threshold: 100%
      status: "exhausted"
      actions:
        - "Incident declared automatically"
        - "Feature releases blocked"
        - "All hands focus on reliability"
        - "Executive escalation"

    - threshold: 80%
      status: "warning"
      actions:
        - "Increase monitoring frequency"
        - "Review pending deployments for risk"
        - "Defer non-critical releases"
        - "Daily reliability review"

    - threshold: 50%
      status: "caution"
      actions:
        - "Heightened awareness"
        - "Additional testing for releases"

    - threshold: 0%
      status: "healthy"
      actions:
        - "Normal operations"
        - "Consider investing in new features"

Decision Framework

> 50% Remaining20-50% Remaining0-20% RemainingExhaustedError Budget CheckBudget Status?HealthyCautionWarningCriticalNormal deploymentvelocityConsider featureworkExtra testingrequiredReview deploymentriskReduce deploymentfrequencyFocus on reliabilityDaily reviewsBlock featurereleasesAll hands onreliabilityIncident declared> 50% Remaining20-50% Remaining0-20% RemainingExhaustedError Budget CheckBudget Status?HealthyCautionWarningCriticalNormal deploymentvelocityConsider featureworkExtra testingrequiredReview deploymentriskReduce deploymentfrequencyFocus on reliabilityDaily reviewsBlock featurereleasesAll hands onreliabilityIncident declared

Burn Rate Alerts

Burn Rate Concepts

Burn rate measures how quickly you're consuming error budget. A burn rate of 1 means you're consuming budget at exactly the rate that would exhaust it by the end of the window.

Burn Rate = (Current Error Rate) / (Maximum Allowed Error Rate)

Example:
SLO: 99.9% (0.1% error budget)
Current error rate: 1% (10× the budget)
Burn rate: 1% / 0.1% = 10

At this rate, error budget exhausted in: 30 days / 10 = 3 days

Multi-Window, Multi-Burn-Rate Alerting

Error BudgetMonitoringFast Burn AlertModerate Burn AlertSlow Burn Alert1-hour windowBurn rate > 14.4Budget depleted in&lt; 2 daysPage immediately6-hour windowBurn rate > 6Budget depleted in&lt; 5 daysAlert duringbusiness hours3-day windowBurn rate > 1Budget depleted byend of monthTicket forinvestigationError BudgetMonitoringFast Burn AlertModerate Burn AlertSlow Burn Alert1-hour windowBurn rate > 14.4Budget depleted in&lt; 2 daysPage immediately6-hour windowBurn rate > 6Budget depleted in&lt; 5 daysAlert duringbusiness hours3-day windowBurn rate > 1Budget depleted byend of monthTicket forinvestigation

Alert Configuration Examples

Fast Burn (Page):

# High severity: Budget exhausted in < 2 days
alert: SLOBurnRateFast
expr: |
  (
    sum(rate(http_requests_total{status=~"5.."}[1h]))
    /
    sum(rate(http_requests_total[1h]))
  ) > (14.4 * (1 - 0.999))
  and
  (
    sum(rate(http_requests_total{status=~"5.."}[5m]))
    /
    sum(rate(http_requests_total[5m]))
  ) > (14.4 * (1 - 0.999))
for: 2m
severity: page
labels:
  burn_rate: "fast"
  depletion_time: "2d"
annotations:
  summary: "Fast error budget burn: exhaustion in < 2 days"

Moderate Burn (Alert):

# Medium severity: Budget exhausted in < 5 days
alert: SLOBurnRateModerate
expr: |
  (
    sum(rate(http_requests_total{status=~"5.."}[6h]))
    /
    sum(rate(http_requests_total[6h]))
  ) > (6 * (1 - 0.999))
  and
  (
    sum(rate(http_requests_total{status=~"5.."}[30m]))
    /
    sum(rate(http_requests_total[30m]))
  ) > (6 * (1 - 0.999))
for: 15m
severity: alert
labels:
  burn_rate: "moderate"
  depletion_time: "5d"

Slow Burn (Ticket):

# Low severity: Budget exhausted by end of window
alert: SLOBurnRateSlow
expr: |
  (
    sum(rate(http_requests_total{status=~"5.."}[3d]))
    /
    sum(rate(http_requests_total[3d]))
  ) > (1 * (1 - 0.999))
  and
  (
    sum(rate(http_requests_total{status=~"5.."}[6h]))
    /
    sum(rate(http_requests_total[6h]))
  ) > (1 * (1 - 0.999))
for: 1h
severity: ticket
labels:
  burn_rate: "slow"
  depletion_time: "30d"

Burn Rate Thresholds

For a 30-day SLO window:

Alert Type Window Burn Rate Time to Exhaustion Action
Critical 1 hour 14.4× < 2 days Page on-call
High 6 hours 6× < 5 days Alert team
Medium 1 day 3× < 10 days Create ticket
Low 3 days 1× 30 days Monitor

Calculation for burn rate thresholds:

Burn Rate = (Window Duration) / (Acceptable Depletion Time)

For 2-day depletion on 30-day window:
Burn Rate = 30 / 2 = 15 (rounded to 14.4 for 5% budget consumption)

Instrumentation

Availability Instrumentation

Application-level tracking:

# Python Flask example with Prometheus
from flask import Flask, request
from prometheus_client import Counter, Histogram, generate_latest
import time

app = Flask(__name__)

# Metrics
request_count = Counter(
    'http_requests_total',
    'Total HTTP requests',
    ['method', 'endpoint', 'status']
)

request_latency = Histogram(
    'http_request_duration_seconds',
    'HTTP request latency',
    ['method', 'endpoint']
)

@app.before_request
def before_request():
    request.start_time = time.time()

@app.after_request
def after_request(response):
    # Record request count
    request_count.labels(
        method=request.method,
        endpoint=request.endpoint or 'unknown',
        status=response.status_code
    ).inc()

    # Record latency
    if hasattr(request, 'start_time'):
        duration = time.time() - request.start_time
        request_latency.labels(
            method=request.method,
            endpoint=request.endpoint or 'unknown'
        ).observe(duration)

    return response

@app.route('/api/users')
def get_users():
    # Your application logic
    return {'users': []}

@app.route('/metrics')
def metrics():
    return generate_latest()

Nginx instrumentation:

# nginx.conf with request tracking
http {
    log_format sli_format '$remote_addr - $remote_user [$time_local] '
                          '"$request" $status $body_bytes_sent '
                          '$request_time $upstream_response_time '
                          '"$http_user_agent"';

    access_log /var/log/nginx/access.log sli_format;

    # Export metrics to Prometheus
    server {
        location /metrics {
            stub_status on;
            access_log off;
        }
    }
}

Latency Instrumentation

Using Prometheus histograms:

# Python: Latency tracking with custom buckets
from prometheus_client import Histogram

# Define buckets relevant to your SLOs
http_latency = Histogram(
    'http_request_duration_seconds',
    'HTTP request latency in seconds',
    ['method', 'endpoint'],
    buckets=[0.1, 0.3, 0.5, 1.0, 2.0, 5.0, 10.0]  # SLO-aligned buckets
)

# Usage
with http_latency.labels(method='GET', endpoint='/api/users').time():
    # Your application logic
    process_request()

Node.js Express example:

// Express middleware for SLI tracking
const promClient = require('prom-client');

const httpRequestDuration = new promClient.Histogram({
  name: 'http_request_duration_seconds',
  help: 'Duration of HTTP requests in seconds',
  labelNames: ['method', 'route', 'status_code'],
  buckets: [0.1, 0.3, 0.5, 1.0, 2.0, 5.0]
});

const httpRequestTotal = new promClient.Counter({
  name: 'http_requests_total',
  help: 'Total number of HTTP requests',
  labelNames: ['method', 'route', 'status_code']
});

app.use((req, res, next) => {
  const start = Date.now();

  res.on('finish', () => {
    const duration = (Date.now() - start) / 1000;

    httpRequestDuration.labels(
      req.method,
      req.route?.path || 'unknown',
      res.statusCode
    ).observe(duration);

    httpRequestTotal.labels(
      req.method,
      req.route?.path || 'unknown',
      res.statusCode
    ).inc();
  });

  next();
});

Prometheus Recording Rules

Pre-aggregate SLI calculations for efficient querying:

# prometheus_rules.yml
groups:
  - name: sli_rules
    interval: 30s
    rules:
      # Availability SLI
      - record: sli:availability:ratio_rate5m
        expr: |
          sum(rate(http_requests_total{status=~"2..|304|401|402|403"}[5m]))
          /
          sum(rate(http_requests_total[5m]))

      # Latency SLI (proportion fast)
      - record: sli:latency:ratio_rate5m
        expr: |
          sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
          /
          sum(rate(http_request_duration_seconds_count[5m]))

      # Error budget consumption (30-day)
      - record: sli:error_budget:consumed_ratio_30d
        expr: |
          1 - (
            (1 - sli:availability:ratio_rate5m)
            /
            (1 - 0.999)  # SLO target
          )

Reporting and Review

SLO Dashboard Components

Key metrics to display:

  1. Current SLI value (real-time)
  2. SLO target (static reference)
  3. Error budget remaining (percentage and time)
  4. Burn rate (current rate)
  5. Trend graph (30-day history)
  6. Time to exhaustion (at current burn rate)

Grafana dashboard JSON snippet:

{
  "panels": [
    {
      "title": "Availability SLI",
      "targets": [
        {
          "expr": "sli:availability:ratio_rate5m * 100"
        }
      ],
      "thresholds": [
        {"value": 99.9, "color": "red"},
        {"value": 99.95, "color": "yellow"},
        {"value": 100, "color": "green"}
      ]
    },
    {
      "title": "Error Budget Remaining",
      "targets": [
        {
          "expr": "sli:error_budget:consumed_ratio_30d * 100"
        }
      ],
      "gauge": {
        "maxValue": 100,
        "minValue": 0,
        "thresholds": [
          {"value": 0, "color": "red"},
          {"value": 20, "color": "orange"},
          {"value": 50, "color": "yellow"},
          {"value": 80, "color": "green"}
        ]
      }
    }
  ]
}

Review Cadences

# Recommended review schedule
slo_review_schedule:

  daily:
    - audience: "Engineering team"
    - review: "Error budget status"
    - action: "Adjust deployment velocity if needed"
    - duration: "5 minutes in standup"

  weekly:
    - audience: "Service owners + SRE"
    - review: "SLO compliance, trends, incidents"
    - action: "Identify reliability improvements"
    - duration: "30 minutes"

  monthly:
    - audience: "Engineering + Product + Leadership"
    - review: "SLO performance, error budget usage, policy effectiveness"
    - action: "Adjust SLOs if needed, prioritise reliability work"
    - duration: "1 hour"

  quarterly:
    - audience: "All stakeholders"
    - review: "SLO strategy, targets, new services"
    - action: "Set reliability goals for next quarter"
    - duration: "2 hours"

Review Meeting Template

# Weekly SLO Review - [Date]

## Services Reviewed
- Service A (payment-api)
- Service B (user-service)
- Service C (notification-service)

## SLO Compliance

| Service | SLO Target | Current SLI | Status | Error Budget |
|---------|-----------|-------------|--------|--------------|
| payment-api | 99.9% | 99.92% | ✅ Healthy | 20% remaining |
| user-service | 99.95% | 99.89% | ⚠️ Warning | -120% (exhausted) |
| notification-service | 99.5% | 99.87% | ✅ Healthy | 74% remaining |

## Incidents Impact
- INC-1234: Database failover (user-service) - consumed 80% of monthly budget
- INC-1235: API timeout spike (payment-api) - consumed 15% of monthly budget

## Actions
1. user-service: Release freeze until error budget recovers
2. user-service: Root cause analysis scheduled for tomorrow
3. payment-api: Continue monitoring, normal operations

## Reliability Improvements
- Implement circuit breaker for user-service database calls
- Add retry logic with exponential backoff for payment-api

Complete Examples

Example 1: E-commerce Checkout SLO

service: checkout-api
description: "Payment processing and order completion"

slis:
  - name: availability
    description: "Proportion of successful checkout requests"
    measurement: |
      sum(rate(http_requests_total{service="checkout",status=~"2.."}[5m]))
      /
      sum(rate(http_requests_total{service="checkout"}[5m]))
    valid_events: "All POST /api/checkout requests"
    good_events: "HTTP 200-299 responses"

  - name: latency
    description: "Proportion of fast checkout requests"
    measurement: |
      sum(rate(http_request_duration_seconds_bucket{service="checkout",le="2.0"}[5m]))
      /
      sum(rate(http_request_duration_seconds_count{service="checkout"}[5m]))
    valid_events: "All POST /api/checkout requests"
    good_events: "Requests completed in < 2 seconds"

slos:
  - sli: availability
    target: 99.95%
    window: 30d

  - sli: latency
    target: 99.0%
    window: 30d

error_budget_policy:
  - remaining: 100%
    actions: ["Deployment freeze", "Incident declared", "All hands"]
  - remaining: 80%
    actions: ["Reduce deployment frequency", "Extra testing"]
  - remaining: 50%
    actions: ["Heightened monitoring"]

alerts:
  - name: CheckoutAvailabilityBurnFast
    expr: |
      (
        1 - sum(rate(http_requests_total{service="checkout",status=~"2.."}[1h]))
            / sum(rate(http_requests_total{service="checkout"}[1h]))
      ) > (14.4 * 0.0005)
    severity: page

  - name: CheckoutLatencyBurnModerate
    expr: |
      (
        1 - sum(rate(http_request_duration_seconds_bucket{service="checkout",le="2.0"}[6h]))
            / sum(rate(http_request_duration_seconds_count{service="checkout"}[6h]))
      ) > (6 * 0.01)
    severity: alert

Example 2: Data Pipeline SLO

service: data-ingestion-pipeline
description: "Real-time event processing pipeline"

slis:
  - name: freshness
    description: "Proportion of events processed within SLA"
    measurement: |
      sum(rate(events_processed{lag_seconds<300}[5m]))
      /
      sum(rate(events_total[5m]))
    valid_events: "All events received"
    good_events: "Events processed within 5 minutes"

  - name: correctness
    description: "Proportion of events processed without errors"
    measurement: |
      sum(rate(events_processed{status="success"}[5m]))
      /
      sum(rate(events_processed[5m]))
    valid_events: "All processing attempts"
    good_events: "Successfully processed events"

slos:
  - sli: freshness
    target: 99.0%
    window: 7d

  - sli: correctness
    target: 99.9%
    window: 7d

error_budget_policy:
  - remaining: 100%
    actions: ["Pause data source integration", "Investigate immediately"]
  - remaining: 50%
    actions: ["Monitor closely", "Prepare rollback plan"]

Quick Reference

SLO Target Selection

Service Criticality Availability Latency (p95) Window
Critical user-facing 99.95% - 99.99% 100-300ms 28-30d
Standard user-facing 99.5% - 99.9% 300-1000ms 28-30d
Internal service 99% - 99.5% 1000-3000ms 7-28d
Batch/Background 95% - 99% N/A 7-28d

Error Budget Time Allowances

SLO Target Downtime/30 days Downtime/year
99.99% 4.32 minutes 52.6 minutes
99.95% 21.6 minutes 4.38 hours
99.9% 43.2 minutes 8.76 hours
99.5% 3.6 hours 1.83 days
99% 7.2 hours 3.65 days
95% 36 hours 18.25 days

Common Prometheus Queries

# Availability SLI (5-minute window)
sum(rate(http_requests_total{status=~"2..|304"}[5m]))
/
sum(rate(http_requests_total[5m]))

# Latency SLI (proportion fast, 5-minute window)
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))

# Error budget remaining (30-day window, 99.9% target)
1 - (
  (1 - sli:availability:ratio_rate5m)
  /
  (1 - 0.999)
)

# Burn rate (1-hour window)
(
  sum(rate(http_requests_total{status=~"5.."}[1h]))
  /
  sum(rate(http_requests_total[1h]))
)
/
(1 - 0.999)

# Time to error budget exhaustion (hours)
(
  sli:error_budget:consumed_ratio_30d * (30 * 24)
)
/
(
  sum(rate(http_requests_total{status=~"5.."}[1h]))
  /
  sum(rate(http_requests_total[1h]))
)

Burn Rate Alert Thresholds (30-day window)

Window Burn Rate Exhaustion Time Severity
1 hour 14.4× 2 days Page
6 hours 6× 5 days Alert
1 day 3× 10 days Warn
3 days 1× 30 days Info

Error Budget Policy Actions

Budget Remaining Status Actions
> 50% Healthy Normal velocity, consider feature work
20-50% Caution Extra testing, review risky changes
0-20% Warning Reduce deployments, daily reviews, focus on reliability
Exhausted Critical Freeze features, all hands on reliability, incident declared

Common Issues and Solutions

Issue: SLI doesn't reflect user experience

Symptoms:

  • Users report problems but SLI shows green
  • SLI is 100% but users experience errors

Solutions:

# Add client-side measurements
- Implement Real User Monitoring (RUM)
- Track synthetic probes from user locations
- Include mobile app metrics
- Monitor from outside your network

# Example: Add synthetic monitoring
apiVersion: v1
kind: ConfigMap
metadata:
  name: blackbox-exporter
data:
  config.yml: |
    modules:
      http_2xx:
        prober: http
        timeout: 5s
        http:
          valid_status_codes: [200, 201, 204]
          fail_if_not_ssl: true
          preferred_ip_protocol: "ip4"

Issue: Error budget exhausted too quickly

Symptoms:

  • Frequently hitting error budget limits
  • Constant deployment freezes
  • Team velocity severely impacted

Solutions:

  1. Re-evaluate SLO target - May be too aggressive

    Current: 99.99% (4.32 min/month)
    Proposed: 99.95% (21.6 min/month) - 5× more budget
    
  2. Exclude expected failures from SLI

    # Don't count 4xx client errors (except 429) in availability
    sum(rate(http_requests_total{status=~"2..|304|4.."}[5m]))
    
  3. Implement gradual rollouts to reduce blast radius

    deployment_strategy:
      - canary: 5%
        duration: 30m
      - canary: 25%
        duration: 1h
      - canary: 100%
    

Issue: Alert fatigue from burn rate alerts

Symptoms:

  • Too many burn rate alerts
  • Alerts during normal operations
  • Team ignoring alerts

Solutions:

# Adjust burn rate thresholds
# Before: Alert at 6× burn rate (too sensitive)
# After: Alert at 10× burn rate

# Add multi-window confirmation
- expr: |
    (error_rate[1h] > threshold)  # Short window
    and
    (error_rate[6h] > threshold)  # Long window confirmation
  for: 15m  # Must persist for 15 minutes

# Reduce noise during deployments
- expr: |
    slo_burn_rate > 10
    and
    absent(deployment_in_progress{service="api"})

Issue: Cannot meet SLO with dependencies

Symptoms:

  • Your SLO: 99.9%
  • Dependency SLO: 99.5%
  • Can't achieve target

Solutions:

# 1. Adjust your SLO based on dependency chain
#    If you depend on 3 services at 99.5% each:
#    Maximum achievable: 99.5% × 99.5% × 99.5% = 98.5%
#    Set your SLO realistically: 98.0%

# 2. Implement resilience patterns
patterns:
  - circuit_breaker:
      failure_threshold: 5
      timeout: 30s
      recovery_time: 60s

  - retry_with_backoff:
      max_attempts: 3
      initial_delay: 100ms
      multiplier: 2

  - fallback:
      cache: stale_data
      timeout: 5s

  - timeout:
      connection: 1s
      request: 5s

# 3. Cache aggressively
cache_strategy:
  ttl: 300s
  stale_while_revalidate: 600s
  serve_stale_on_error: true

Issue: Latency SLI gaming

Symptoms:

  • System terminates slow requests to meet SLO
  • Users see more failures but latency SLI looks good

Solutions:

# Use both availability AND latency SLOs
slos:
  - name: availability
    target: 99.9%  # Can't just drop requests

  - name: latency
    target: 99.0%  # Must be fast AND successful

# Track all outcomes
sli_definition:
  valid_events: "All requests initiated by users"
  good_events: "Requests that completed successfully AND within latency target"

  # Bad approach: Excludes timeouts
  # good_events: "Successful requests < 300ms"

  # Good approach: Includes all outcomes
  # Timeout counts as both availability AND latency failure

Issue: Different user expectations by region

Symptoms:

  • Global SLO doesn't reflect regional experience
  • Some regions consistently poor

Solutions:

# Define SLOs per region
slos:
  - region: us-east
    availability: 99.95%
    latency_p95: 100ms

  - region: eu-west
    availability: 99.95%
    latency_p95: 150ms

  - region: ap-southeast
    availability: 99.9%
    latency_p95: 300ms

# Prometheus query with regional labels
sum(rate(http_requests_total{status=~"2..",region="us-east"}[5m]))
/
sum(rate(http_requests_total{region="us-east"}[5m]))

Issue: Monthly SLO resets create perverse incentives

Symptoms:

  • Team "saves" error budget for end of month
  • Riskier deploys at month start
  • Sudden focus on reliability at month end

Solutions:

# Use rolling windows instead of calendar windows
slo:
  window: 28d  # Rolling 28-day window
  target: 99.9%

# Or use multiple windows
slos:
  - window: 7d   # Short-term quality signal
    target: 99.5%

  - window: 28d  # Long-term trend
    target: 99.9%

# Both must be met

Issue: Can't identify which component caused SLO violation

Symptoms:

  • SLO alert fires but unclear which service/component is responsible
  • Long investigation time

Solutions:

# Add detailed labels to track attribution
http_requests_total{
  service="api",
  component="database",
  operation="query",
  failure_mode="timeout"
}

# Create SLIs per critical component
- sli:availability:database
- sli:availability:cache
- sli:availability:external_api

# Use distributed tracing
# Tag requests with trace_id and analyse:
- Which service added the most latency?
- Which component had the error?
- What was the failure propagation path?