Chaos Engineering Practices
A comprehensive guide to implementing chaos engineering for building resilient systems through controlled failure injection and experimentation.
Overview
Chaos Engineering is the discipline of experimenting on a system to build confidence in its capability to withstand turbulent conditions in production. Rather than waiting for failures to occur naturally, teams proactively inject faults to discover weaknesses before they impact users.
flowchart TB
subgraph "Chaos Engineering Cycle"
A[Define Steady State] --> B[Formulate Hypothesis]
B --> C[Design Experiment]
C --> D[Define Blast Radius]
D --> E[Set Safeguards]
E --> F[Run Experiment]
F --> G{Hypothesis<br/>Validated?}
G -->|Yes| H[Expand Scope]
G -->|No| I[Fix Weakness]
I --> J[Document Learnings]
H --> J
J --> A
end
subgraph "Continuous Validation"
F -.-> K[Monitor Observability]
K -.-> L{Health Check<br/>Failing?}
L -->|Yes| M[Abort & Rollback]
L -->|No| F
M --> I
end
style A fill:#e1f5fe
style F fill:#fff3e0
style I fill:#fce4ec
style M fill:#ffcdd2
Core Principles:
- Hypothesise about steady state - Define normal system behaviour
- Vary real-world events - Inject realistic failure scenarios
- Run experiments in production - Test where it matters most
- Automate experiments - Continuous chaos for continuous verification
- Minimise blast radius - Start small, expand gradually
Defining Steady State and Hypotheses
Key Concepts
Steady State represents normal system behaviour measured through business and technical metrics. It should focus on user-observable outcomes rather than internal system attributes.
Hypothesis Format:
Given [normal system state]
When [specific failure is injected]
Then [expected system behaviour]
And [user impact remains within acceptable bounds]
Steady State Metrics
| Metric Type | Examples | Purpose |
|---|---|---|
| Business Metrics | Orders/second, conversion rate, revenue | Measure actual user impact |
| Application Metrics | Request latency (p50/p99), error rate, throughput | Track service health |
| Infrastructure Metrics | CPU/memory usage, network bandwidth, disk I/O | Identify resource constraints |
| User Experience | Page load time, transaction success rate | Direct impact measurement |
Example Hypotheses
# Hypothesis 1: Service mesh retry resilience
steady_state:
metric: "http_request_success_rate"
threshold: "> 99.5%"
duration: "5m"
hypothesis: |
When 30% of payment-service pods fail,
Then Istio retry logic will maintain > 99.5% success rate,
And p99 latency will increase by < 50ms,
And no user transactions will be lost.
validation_metrics:
- name: "success_rate"
query: "sum(rate(http_requests_total{status!~'5..'}[5m])) / sum(rate(http_requests_total[5m]))"
threshold: "> 0.995"
- name: "p99_latency"
query: "histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))"
threshold: "< 0.250" # 200ms baseline + 50ms tolerance
# Hypothesis 2: Database connection pool exhaustion
steady_state:
metric: "database_query_success_rate"
threshold: "> 99.9%"
duration: "10m"
hypothesis: |
When database connection pool is exhausted (100 connections),
Then application gracefully degrades with circuit breaker,
And users see cached data rather than errors,
And system auto-recovers when connections available.
validation_metrics:
- name: "circuit_breaker_state"
query: "circuit_breaker_open{service='api'}"
threshold: "== 1" # Should open
- name: "cache_hit_rate"
query: "rate(cache_hits[5m]) / rate(cache_requests[5m])"
threshold: "> 0.80"
Steady State Definition Template
## Steady State Definition: [Service Name]
### Business Metrics
- **Orders Processed:** 1,200/hour (±10%)
- **Conversion Rate:** 3.5% (±0.5%)
- **Average Order Value:** £45 (±£5)
### Technical Metrics
- **Request Rate:** 500 req/s (±50 req/s)
- **Error Rate:** < 0.1% (5xx responses)
- **P50 Latency:** < 100ms
- **P99 Latency:** < 300ms
- **CPU Utilisation:** 40-60%
- **Memory Usage:** 2-4 GB
### Dependencies
- **Database:** Read latency < 10ms, Write latency < 25ms
- **Cache:** Hit rate > 85%
- **External API:** Response time < 500ms, Availability > 99.9%
### Acceptable Degradation Bounds
- Latency increase: < 50% from baseline
- Throughput decrease: < 20% from baseline
- Error rate: Must remain < 1%
Blast Radius and Scope Control
Key Concepts
Blast Radius defines the scope and potential impact of a chaos experiment. Starting with minimal blast radius and gradually expanding is critical for safe chaos engineering in production.
flowchart LR
subgraph "Blast Radius Progression"
A[Single Pod<br/>Dev Environment] --> B[Pod Group<br/>Staging]
B --> C[Single AZ<br/>Staging]
C --> D[Single Pod<br/>Production]
D --> E[10% Traffic<br/>Production]
E --> F[Single AZ<br/>Production]
F --> G[Region<br/>Production]
end
style A fill:#c8e6c9
style D fill:#fff9c4
style G fill:#ffccbc
Blast Radius Dimensions
| Dimension | Controls | Examples |
|---|---|---|
| Environment | Where experiment runs | Dev → Staging → Production |
| Infrastructure | Resources affected | Single pod → Node → AZ → Region |
| Traffic | Percentage of requests | 1% → 10% → 50% → 100% |
| Time | Duration of experiment | 30s → 5m → 30m → Continuous |
| User Scope | User population | Internal → Beta users → All users |
Blast Radius Configuration Examples
# Chaos Mesh - Pod failure with controlled scope
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
name: payment-service-failure
namespace: chaos-testing
spec:
action: pod-failure
mode: fixed-percent # Control blast radius
value: "30" # Affect 30% of pods
duration: "2m"
selector:
namespaces:
- production
labelSelectors:
app: payment-service
version: v2.1
# Additional filters to limit scope
expressionSelectors:
- key: chaos.enabled
operator: In
values: ["true"]
# Schedule for controlled timing
scheduler:
cron: "@every 6h"
---
# Litmus - Network latency with gradual ramp-up
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: network-latency-rampup
namespace: production
spec:
engineState: active
appinfo:
appns: production
applabel: app=checkout-service
appkind: deployment
chaosServiceAccount: litmus-admin
experiments:
- name: pod-network-latency
spec:
components:
env:
- name: NETWORK_LATENCY
value: "2000" # 2s latency
- name: JITTER
value: "500" # ±500ms
- name: TARGET_PODS
value: "10%" # Start with 10% blast radius
- name: TOTAL_CHAOS_DURATION
value: "300" # 5 minutes
- name: RAMP_TIME
value: "60" # Gradual ramp-up over 1 minute
Blast Radius Decision Matrix
| Experiment Type | Starting Scope | Production-Ready Scope | Notes |
|---|---|---|---|
| Pod Failure | 1 pod in staging | 10-30% of pods | Ensure replicas > 3 |
| Network Latency | 1% traffic | 25-50% traffic | Monitor p99 latency |
| CPU Stress | Single pod, 50% load | 30% pods, 80% load | Watch for cascade |
| Memory Pressure | 1 pod, 70% memory | 20% pods, 85% memory | Risk of OOM kills |
| DNS Failure | Non-critical service | Critical service, 10% pods | High impact potential |
| Zone Failure | Staging only | Production with 3+ zones | Major blast radius |
Common Chaos Experiments
Pod and Container Failures
# Chaos Mesh - Pod Kill
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
name: pod-kill-experiment
spec:
action: pod-kill
mode: one
selector:
namespaces:
- production
labelSelectors:
app: api-service
scheduler:
cron: "@every 4h"
---
# Chaos Mesh - Container Kill (partial pod failure)
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
name: container-kill-experiment
spec:
action: container-kill
mode: fixed
value: "2"
containerNames:
- sidecar-proxy # Test sidecar resilience
selector:
namespaces:
- production
labelSelectors:
app: payment-service
duration: "5m"
Network Chaos
# Chaos Mesh - Network Partition
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: network-partition
spec:
action: partition
mode: all
selector:
namespaces:
- production
labelSelectors:
app: order-service
direction: to
target:
selector:
namespaces:
- production
labelSelectors:
app: inventory-service
duration: "3m"
---
# Chaos Mesh - Network Delay
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: network-latency-injection
spec:
action: delay
mode: fixed-percent
value: "50"
delay:
latency: "250ms"
correlation: "50" # 50% correlation between packets
jitter: "50ms"
selector:
namespaces:
- production
labelSelectors:
app: checkout-service
duration: "10m"
---
# Chaos Mesh - Packet Loss
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: packet-loss
spec:
action: loss
mode: one
loss:
loss: "25" # 25% packet loss
correlation: "25"
selector:
namespaces:
- production
labelSelectors:
app: notification-service
direction: to
target:
selector:
namespaces:
- production
labelSelectors:
app: email-gateway
duration: "5m"
---
# Chaos Mesh - Bandwidth Limitation
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: bandwidth-limit
spec:
action: bandwidth
mode: all
bandwidth:
rate: "1mbps" # Limit to 1 Mbps
limit: 20000 # Buffer size
buffer: 10000 # Peakrate buffer
selector:
namespaces:
- production
labelSelectors:
app: video-streaming
duration: "15m"
Resource Stress
# Chaos Mesh - CPU Stress
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
name: cpu-stress
spec:
mode: fixed-percent
value: "40"
stressors:
cpu:
workers: 4
load: 80 # 80% CPU load
selector:
namespaces:
- production
labelSelectors:
app: analytics-service
duration: "10m"
---
# Chaos Mesh - Memory Stress
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
name: memory-pressure
spec:
mode: one
stressors:
memory:
workers: 4
size: "512MB" # Allocate 512MB
selector:
namespaces:
- production
labelSelectors:
app: cache-service
duration: "5m"
---
# Litmus - Disk Fill
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: disk-fill
spec:
engineState: active
appinfo:
appns: production
applabel: app=database
appkind: statefulset
chaosServiceAccount: litmus-admin
experiments:
- name: disk-fill
spec:
components:
env:
- name: FILL_PERCENTAGE
value: "80" # Fill disk to 80%
- name: TARGET_CONTAINER
value: "postgres"
- name: TOTAL_CHAOS_DURATION
value: "300"
Application-Level Failures
# Chaos Mesh - HTTP Abort (Service unavailable)
apiVersion: chaos-mesh.org/v1alpha1
kind: HTTPChaos
metadata:
name: http-abort
spec:
mode: fixed-percent
value: "25"
target: Request
port: 8080
method: POST
path: /api/v1/orders
abort: true
selector:
namespaces:
- production
labelSelectors:
app: order-service
duration: "5m"
---
# Chaos Mesh - HTTP Delay
apiVersion: chaos-mesh.org/v1alpha1
kind: HTTPChaos
metadata:
name: http-latency
spec:
mode: all
target: Request
port: 8080
path: /api/v1/payments
delay: "3s"
selector:
namespaces:
- production
labelSelectors:
app: payment-gateway
duration: "10m"
---
# Gremlin - DNS Failure
gremlin attack create dns \
--hostname database.internal.example.com \
--length 300 \
--tags "service:api,env:production" \
--target-type container \
--target-percent 20
Infrastructure Failures
# Chaos Mesh - Kernel I/O Chaos
apiVersion: chaos-mesh.org/v1alpha1
kind: IOChaos
metadata:
name: io-latency
spec:
action: latency
mode: one
volumePath: /var/lib/postgresql/data
path: /var/lib/postgresql/data/**/*
delay: "100ms"
percent: 50
selector:
namespaces:
- production
labelSelectors:
app: postgres
duration: "10m"
---
# Litmus - Node Drain (simulates node failure)
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: node-drain-chaos
spec:
engineState: active
chaosServiceAccount: litmus-admin
experiments:
- name: node-drain
spec:
components:
env:
- name: TARGET_NODE
value: "ip-10-0-1-50.eu-west-1.compute.internal"
- name: TOTAL_CHAOS_DURATION
value: "600" # 10 minutes
---
# AWS Fault Injection Simulator - AZ outage
{
"description": "Simulate AZ failure",
"targets": {
"Subnets": {
"resourceType": "aws:ec2:subnet",
"resourceArns": [
"arn:aws:ec2:eu-west-1:123456789012:subnet/subnet-abc123"
],
"selectionMode": "ALL"
}
},
"actions": {
"BlockSubnet": {
"actionId": "aws:network:disrupt-connectivity",
"parameters": {
"duration": "PT5M",
"scope": "all"
},
"targets": {
"Subnets": "Subnets"
}
}
},
"stopConditions": [
{
"source": "aws:cloudwatch:alarm",
"value": "arn:aws:cloudwatch:eu-west-1:123456789012:alarm:HighErrorRate"
}
],
"roleArn": "arn:aws:iam::123456789012:role/FISRole"
}
Chaos Engineering Tools
Chaos Mesh (Kubernetes-native)
# Install Chaos Mesh
helm repo add chaos-mesh https://charts.chaos-mesh.org
helm install chaos-mesh chaos-mesh/chaos-mesh \
--namespace=chaos-mesh \
--create-namespace \
--set dashboard.create=true
# Access dashboard
kubectl port-forward -n chaos-mesh svc/chaos-dashboard 2333:2333
# Create experiment via CLI
kubectl apply -f - <<EOF
apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
name: payment-resilience-test
namespace: chaos-testing
spec:
entry: entry
templates:
- name: entry
templateType: Serial
children:
- network-delay
- pod-failure
- validate-recovery
- name: network-delay
templateType: NetworkChaos
deadline: 5m
networkChaos:
action: delay
mode: fixed-percent
value: "30"
delay:
latency: "500ms"
selector:
namespaces: [production]
labelSelectors:
app: payment-service
- name: pod-failure
templateType: PodChaos
deadline: 3m
podChaos:
action: pod-kill
mode: fixed
value: "2"
selector:
namespaces: [production]
labelSelectors:
app: payment-service
- name: validate-recovery
templateType: Suspend
deadline: 10m
EOF
# List experiments
kubectl get podchaos,networkchaos,stresschaos -A
# Pause experiment
kubectl annotate podchaos pod-kill-experiment experiment.chaos-mesh.org/pause=true
# Resume experiment
kubectl annotate podchaos pod-kill-experiment experiment.chaos-mesh.org/pause-
# Delete experiment
kubectl delete podchaos pod-kill-experiment
Litmus Chaos
# Install Litmus
kubectl apply -f https://litmuschaos.github.io/litmus/litmus-operator-latest.yaml
# Install chaos experiments
kubectl apply -f https://hub.litmuschaos.io/api/chaos/master?file=charts/generic/experiments.yaml
# Create service account
kubectl apply -f - <<EOF
apiVersion: v1
kind: ServiceAccount
metadata:
name: litmus-admin
namespace: production
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: litmus-admin
rules:
- apiGroups: [""]
resources: ["pods", "events"]
verbs: ["create", "delete", "get", "list", "patch", "update"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets"]
verbs: ["get", "list", "update"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: litmus-admin
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: litmus-admin
subjects:
- kind: ServiceAccount
name: litmus-admin
namespace: production
EOF
# Run experiment
kubectl apply -f chaos-engine.yaml
# Monitor experiment
kubectl describe chaosengine network-latency-rampup -n production
# View results
kubectl logs -n production \
$(kubectl get pods -n production -l chaosUID -o jsonpath='{.items[0].metadata.name}') \
-c chaos-runner
# Abort experiment
kubectl patch chaosengine network-latency-rampup -n production \
--type merge -p '{"spec":{"engineState":"stop"}}'
Gremlin (SaaS Platform)
# Install Gremlin agent
helm repo add gremlin https://helm.gremlin.com
helm install gremlin gremlin/gremlin \
--namespace gremlin \
--create-namespace \
--set gremlin.teamID=$GREMLIN_TEAM_ID \
--set gremlin.teamSecret=$GREMLIN_TEAM_SECRET \
--set gremlin.clusterID=$CLUSTER_NAME
# Create CPU stress attack (via CLI)
gremlin attack create cpu \
--cores 2 \
--percent 80 \
--length 300 \
--tags "service:payment,env:production" \
--target-type container \
--target-percent 25
# Create network latency attack
gremlin attack create latency \
--latency 500 \
--length 600 \
--tags "service:checkout" \
--target-type container
# Create blackhole (network partition)
gremlin attack create blackhole \
--hostname database.internal \
--port 5432 \
--length 180 \
--tags "service:api"
# List active attacks
gremlin attack list --active
# Halt attack
gremlin attack halt <attack-id>
# Create scenario (multi-step experiment)
gremlin scenario create \
--name "Database Failover Test" \
--schedule "0 2 * * MON" \
--steps '[
{
"type": "latency",
"target": {"tags": ["role:primary-db"]},
"args": {"latency": 1000, "length": 300}
},
{
"type": "shutdown",
"target": {"tags": ["role:primary-db"]},
"args": {"delay": 10}
}
]'
AWS Fault Injection Simulator (FIS)
# Create experiment template
aws fis create-experiment-template \
--cli-input-json file://fis-template.json
# Start experiment
EXPERIMENT_ID=$(aws fis start-experiment \
--experiment-template-id EXT123456 \
--query 'experiment.id' \
--output text)
# Monitor experiment
aws fis get-experiment --id $EXPERIMENT_ID
# Stop experiment
aws fis stop-experiment --id $EXPERIMENT_ID
# Example: EC2 instance termination
{
"description": "Test auto-scaling response",
"targets": {
"Instances": {
"resourceType": "aws:ec2:instance",
"resourceTags": {
"Environment": "production",
"ChaosReady": "true"
},
"filters": [
{
"path": "State.Name",
"values": ["running"]
}
],
"selectionMode": "COUNT(2)"
}
},
"actions": {
"TerminateInstances": {
"actionId": "aws:ec2:terminate-instances",
"targets": {
"Instances": "Instances"
}
}
},
"stopConditions": [
{
"source": "aws:cloudwatch:alarm",
"value": "arn:aws:cloudwatch:eu-west-1:123456789012:alarm:UnhealthyHosts"
}
],
"roleArn": "arn:aws:iam::123456789012:role/FISRole"
}
Tool Comparison
| Feature | Chaos Mesh | Litmus | Gremlin | AWS FIS |
|---|---|---|---|---|
| Platform | Kubernetes | Kubernetes | Multi-platform | AWS only |
| Cost | Free | Free | Commercial | Pay-per-use |
| Deployment | Self-hosted | Self-hosted | SaaS + Agent | Managed |
| Experiment Types | 10+ types | 30+ experiments | 15+ attacks | AWS resources |
| GUI | Dashboard | ChaosCenter | Web platform | AWS Console |
| Workflows | Native | Native | Scenarios | Templates |
| Observability | Prometheus metrics | Built-in | Integrated | CloudWatch |
| RBAC | Kubernetes RBAC | Kubernetes RBAC | Built-in | IAM |
| Best For | K8s chaos | K8s + GitOps | Enterprise, compliance | AWS infrastructure |
Safeguards and Abort Conditions
Key Concepts
Safeguards prevent chaos experiments from causing unacceptable harm. Every experiment must have automated abort conditions and manual stop mechanisms.
flowchart TB
A[Start Experiment] --> B{Pre-flight<br/>Checks}
B -->|Pass| C[Inject Fault]
B -->|Fail| D[Abort - Prerequisites Not Met]
C --> E{Monitor<br/>Health Checks}
E -->|Healthy| F{Abort<br/>Condition<br/>Triggered?}
E -->|Unhealthy| G[Abort - Health Check Failed]
F -->|No| H{Duration<br/>Complete?}
F -->|Yes| I[Abort - Condition Met]
H -->|No| E
H -->|Yes| J[Stop Fault Injection]
J --> K{Verify<br/>Recovery}
K -->|Recovered| L[Complete Successfully]
K -->|Not Recovered| M[Alert - Manual Intervention]
D --> N[Cleanup]
G --> N
I --> N
L --> N
M --> N
style D fill:#ffcdd2
style G fill:#ffcdd2
style I fill:#fff9c4
style L fill:#c8e6c9
style M fill:#ffccbc
Pre-flight Checks
# Chaos Mesh - StatusCheck (custom resource)
apiVersion: chaos-mesh.org/v1alpha1
kind: StatusCheck
metadata:
name: preflight-checks
spec:
mode: Synchronous
type: HTTP
intervalSeconds: 5
timeoutSeconds: 10
successThreshold: 3
failureThreshold: 1
http:
url: http://health-check-service/readiness
method: GET
criteria:
statusCode: "200"
# Use in workflow
---
apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
name: safe-experiment
spec:
entry: check-then-chaos
templates:
- name: check-then-chaos
templateType: Serial
children:
- preflight-validation
- run-chaos
- postflight-validation
- name: preflight-validation
templateType: StatusCheck
statusCheck:
mode: Synchronous
type: HTTP
http:
url: http://api-service/health
criteria:
statusCode: "200"
successThreshold: 3
failureThreshold: 1
- name: run-chaos
templateType: PodChaos
# ... chaos spec
Health Check Examples
# Python health check for experiment safety
import requests
import prometheus_client as prom
from typing import Dict, List, Tuple
class ChaosHealthCheck:
def __init__(self, prometheus_url: str):
self.prom_url = prometheus_url
def check_error_rate(self, threshold: float = 0.01) -> Tuple[bool, str]:
"""Verify error rate is below threshold."""
query = """
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
"""
result = self._query_prometheus(query)
if not result:
return False, "Unable to query error rate"
error_rate = result[0]['value'][1]
if float(error_rate) > threshold:
return False, f"Error rate {error_rate} exceeds threshold {threshold}"
return True, "Error rate acceptable"
def check_latency(self, p99_threshold_ms: int = 1000) -> Tuple[bool, str]:
"""Verify p99 latency is below threshold."""
query = f"""
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
) * 1000
"""
result = self._query_prometheus(query)
if not result:
return False, "Unable to query latency"
p99_latency = float(result[0]['value'][1])
if p99_latency > p99_threshold_ms:
return False, f"P99 latency {p99_latency}ms exceeds {p99_threshold_ms}ms"
return True, f"P99 latency {p99_latency}ms acceptable"
def check_service_availability(self, min_replicas: int = 2) -> Tuple[bool, str]:
"""Verify minimum number of healthy replicas."""
query = f"""
sum(up{{job="api-service"}})
"""
result = self._query_prometheus(query)
if not result:
return False, "Unable to query service availability"
available = int(float(result[0]['value'][1]))
if available < min_replicas:
return False, f"Only {available} replicas available, need {min_replicas}"
return True, f"{available} replicas available"
def check_dependency_health(self, dependencies: List[str]) -> Tuple[bool, str]:
"""Verify all dependencies are healthy."""
for dep in dependencies:
query = f'up{{job="{dep}"}}'
result = self._query_prometheus(query)
if not result or float(result[0]['value'][1]) != 1:
return False, f"Dependency {dep} is unhealthy"
return True, "All dependencies healthy"
def run_preflight_checks(self) -> bool:
"""Run all preflight checks before starting experiment."""
checks = [
self.check_error_rate(threshold=0.005), # 0.5% max
self.check_latency(p99_threshold_ms=500),
self.check_service_availability(min_replicas=3),
self.check_dependency_health(['database', 'cache', 'auth-service'])
]
for passed, message in checks:
print(f"{'✓' if passed else '✗'} {message}")
if not passed:
print("❌ Preflight checks failed. Aborting experiment.")
return False
print("✅ All preflight checks passed. Safe to proceed.")
return True
def _query_prometheus(self, query: str) -> list:
"""Execute Prometheus query."""
try:
response = requests.get(
f"{self.prom_url}/api/v1/query",
params={'query': query},
timeout=10
)
return response.json()['data']['result']
except Exception as e:
print(f"Error querying Prometheus: {e}")
return []
# Usage
health_check = ChaosHealthCheck('http://prometheus:9090')
if health_check.run_preflight_checks():
# Start chaos experiment
pass
Abort Conditions
# AWS FIS - CloudWatch Alarm Stop Condition
stopConditions:
- source: aws:cloudwatch:alarm
value: arn:aws:cloudwatch:eu-west-1:123456789012:alarm:HighErrorRate
- source: aws:cloudwatch:alarm
value: arn:aws:cloudwatch:eu-west-1:123456789012:alarm:HighLatency
- source: aws:cloudwatch:alarm
value: arn:aws:cloudwatch:eu-west-1:123456789012:alarm:LowAvailability
---
# Chaos Mesh - StatusCheck for abort
apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
name: chaos-with-abort
spec:
entry: monitored-chaos
templates:
- name: monitored-chaos
templateType: Parallel
children:
- inject-fault
- monitor-health
- name: inject-fault
templateType: PodChaos
deadline: 10m
abortWithStatusCheck: health-monitor # References monitor
podChaos:
action: pod-kill
mode: fixed-percent
value: "20"
selector:
namespaces: [production]
labelSelectors:
app: order-service
- name: monitor-health
templateType: StatusCheck
statusCheck:
mode: Continuous
type: HTTP
intervalSeconds: 10
timeoutSeconds: 5
failureThreshold: 3 # Abort after 3 consecutive failures
http:
url: http://order-service/health
method: GET
criteria:
statusCode: "200"
---
# Custom abort controller in Python
import time
import requests
from datetime import datetime, timedelta
class ChaosAbortController:
def __init__(self, experiment_name: str, prometheus_url: str):
self.experiment_name = experiment_name
self.prom_url = prometheus_url
self.start_time = datetime.now()
self.max_duration = timedelta(minutes=15)
def should_abort(self) -> Tuple[bool, str]:
"""Check if experiment should be aborted."""
# Timeout check
if datetime.now() - self.start_time > self.max_duration:
return True, "Maximum duration exceeded"
# Error rate check
error_rate = self._get_error_rate()
if error_rate and error_rate > 0.05: # 5% threshold
return True, f"Error rate {error_rate:.2%} exceeds 5%"
# Latency check
p99_latency = self._get_p99_latency()
if p99_latency and p99_latency > 2000: # 2s threshold
return True, f"P99 latency {p99_latency}ms exceeds 2000ms"
# Business metric check
throughput = self._get_throughput()
if throughput and throughput < 100: # Minimum req/s
return True, f"Throughput {throughput} req/s below minimum 100 req/s"
# Error budget check
error_budget_remaining = self._get_error_budget_remaining()
if error_budget_remaining is not None and error_budget_remaining < 0.10:
return True, f"Error budget at {error_budget_remaining:.1%}, below 10% threshold"
return False, "All checks passed"
def monitor_and_abort(self, check_interval: int = 30):
"""Continuously monitor and abort if needed."""
while True:
should_abort, reason = self.should_abort()
if should_abort:
print(f"🛑 ABORTING EXPERIMENT: {reason}")
self._abort_experiment()
return
print(f"✓ Health check passed: {reason}")
time.sleep(check_interval)
def _abort_experiment(self):
"""Abort the chaos experiment."""
# Example for Chaos Mesh
import subprocess
subprocess.run([
'kubectl', 'delete', 'podchaos', self.experiment_name,
'--namespace', 'chaos-testing'
])
def _get_error_rate(self) -> float:
query = """
sum(rate(http_requests_total{status=~"5.."}[1m]))
/ sum(rate(http_requests_total[1m]))
"""
return self._query_single_value(query)
def _get_p99_latency(self) -> float:
query = """
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[1m])) by (le)
) * 1000
"""
return self._query_single_value(query)
def _get_throughput(self) -> float:
query = "sum(rate(http_requests_total[1m]))"
return self._query_single_value(query)
def _get_error_budget_remaining(self) -> float:
query = """
1 - (
(1 - (
sum(rate(http_requests_total{status!~"5.."}[30d]))
/ sum(rate(http_requests_total[30d]))
))
/ (1 - 0.999)
)
"""
return self._query_single_value(query)
def _query_single_value(self, query: str) -> float:
try:
response = requests.get(
f"{self.prom_url}/api/v1/query",
params={'query': query},
timeout=5
)
result = response.json()['data']['result']
if result:
return float(result[0]['value'][1])
except:
pass
return None
Circuit Breaker Pattern
# Istio DestinationRule with circuit breaker
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: payment-service-circuit-breaker
spec:
host: payment-service
trafficPolicy:
connectionPool:
tcp:
maxConnections: 100
http:
http1MaxPendingRequests: 50
http2MaxRequests: 100
maxRequestsPerConnection: 2
outlierDetection:
consecutiveErrors: 5
interval: 30s
baseEjectionTime: 30s
maxEjectionPercent: 50
minHealthPercent: 40
# Automatically eject unhealthy pods during chaos
loadBalancer:
simple: LEAST_REQUEST
Observability Validation During Experiments
Key Concepts
Observability is critical for validating chaos experiments. You must monitor system behaviour before, during, and after fault injection to validate hypotheses and detect anomalies.
flowchart LR
subgraph "Baseline Collection"
A[Establish Baseline<br/>T-15min to T0] --> B[Record Normal<br/>Behaviour]
end
subgraph "Experiment Phase"
C[Start Fault Injection<br/>T0] --> D[Monitor Metrics<br/>T0 to T+N]
D --> E[Capture Events<br/>& Logs]
end
subgraph "Recovery Phase"
F[Stop Fault<br/>T+N] --> G[Monitor Recovery<br/>T+N to T+N+30]
G --> H[Verify Steady State<br/>Restored]
end
B --> C
E --> F
subgraph "Analysis"
I[Compare Baseline<br/>vs Experiment] --> J[Validate Hypothesis]
J --> K{Hypothesis<br/>Confirmed?}
K -->|Yes| L[Document Success]
K -->|No| M[Investigate Gap]
end
H --> I
style B fill:#e1f5fe
style E fill:#fff3e0
style H fill:#c8e6c9
style M fill:#ffccbc
Metrics to Monitor
# Prometheus recording rules for chaos experiments
groups:
- name: chaos_experiment_metrics
interval: 10s
rules:
# Request metrics
- record: chaos:request_rate:sum
expr: sum(rate(http_requests_total[1m])) by (service)
- record: chaos:error_rate:ratio
expr: |
sum(rate(http_requests_total{status=~"5.."}[1m])) by (service)
/
sum(rate(http_requests_total[1m])) by (service)
# Latency metrics
- record: chaos:latency:p50
expr: |
histogram_quantile(0.50,
sum(rate(http_request_duration_seconds_bucket[1m])) by (service, le)
)
- record: chaos:latency:p95
expr: |
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[1m])) by (service, le)
)
- record: chaos:latency:p99
expr: |
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[1m])) by (service, le)
)
# Resource metrics
- record: chaos:cpu_usage:avg
expr: avg(rate(container_cpu_usage_seconds_total[1m])) by (pod, container)
- record: chaos:memory_usage:bytes
expr: avg(container_memory_working_set_bytes) by (pod, container)
# Infrastructure metrics
- record: chaos:pod_ready:count
expr: sum(kube_pod_status_ready{condition="true"}) by (namespace, deployment)
- record: chaos:pod_restarts:rate
expr: sum(rate(kube_pod_container_status_restarts_total[5m])) by (namespace, pod)
---
# Grafana dashboard for experiment monitoring
{
"dashboard": {
"title": "Chaos Experiment Dashboard",
"panels": [
{
"title": "Request Rate (Baseline vs Experiment)",
"targets": [
{
"expr": "chaos:request_rate:sum{service='$service'}",
"legendFormat": "Current"
},
{
"expr": "avg_over_time(chaos:request_rate:sum{service='$service'}[15m] offset 30m)",
"legendFormat": "Baseline"
}
]
},
{
"title": "Error Rate",
"targets": [
{
"expr": "chaos:error_rate:ratio{service='$service'} * 100",
"legendFormat": "Error %"
}
],
"alert": {
"conditions": [
{
"evaluator": {
"params": [5],
"type": "gt"
},
"query": {"params": ["A", "5m", "now"]}
}
]
}
},
{
"title": "Latency Percentiles",
"targets": [
{
"expr": "chaos:latency:p50{service='$service'} * 1000",
"legendFormat": "P50"
},
{
"expr": "chaos:latency:p95{service='$service'} * 1000",
"legendFormat": "P95"
},
{
"expr": "chaos:latency:p99{service='$service'} * 1000",
"legendFormat": "P99"
}
]
}
],
"annotations": {
"list": [
{
"name": "Chaos Events",
"datasource": "Prometheus",
"expr": "ALERTS{alertname=~'Chaos.*'}",
"iconColor": "red"
},
{
"name": "Deployments",
"datasource": "Prometheus",
"expr": "changes(kube_deployment_status_observed_generation[30s]) > 0",
"iconColor": "blue"
}
]
}
}
}
Distributed Tracing Validation
# Analyse trace data during chaos experiments
from jaeger_client import Config
from opentelemetry import trace
import time
class ChaosTraceAnalyser:
def __init__(self, jaeger_endpoint: str):
self.jaeger_endpoint = jaeger_endpoint
def analyse_traces_during_experiment(
self,
service_name: str,
start_time: datetime,
end_time: datetime
) -> Dict:
"""Analyse trace patterns during chaos experiment."""
traces = self._fetch_traces(service_name, start_time, end_time)
analysis = {
'total_traces': len(traces),
'error_count': 0,
'retry_count': 0,
'timeout_count': 0,
'circuit_breaker_open': 0,
'latency_spikes': [],
'affected_operations': set()
}
for trace in traces:
# Check for errors
if self._has_error_span(trace):
analysis['error_count'] += 1
analysis['affected_operations'].add(trace['operationName'])
# Detect retries
retry_count = self._count_retry_spans(trace)
if retry_count > 0:
analysis['retry_count'] += retry_count
# Detect timeouts
if self._has_timeout(trace):
analysis['timeout_count'] += 1
# Detect circuit breaker
if self._circuit_breaker_triggered(trace):
analysis['circuit_breaker_open'] += 1
# Identify latency spikes
duration_ms = trace['duration'] / 1000
if duration_ms > 1000: # > 1s
analysis['latency_spikes'].append({
'trace_id': trace['traceID'],
'duration_ms': duration_ms,
'operation': trace['operationName']
})
return analysis
def compare_baseline_vs_experiment(
self,
baseline_analysis: Dict,
experiment_analysis: Dict
) -> Dict:
"""Compare trace patterns between baseline and experiment."""
return {
'error_rate_delta': (
experiment_analysis['error_count'] / experiment_analysis['total_traces']
- baseline_analysis['error_count'] / baseline_analysis['total_traces']
),
'retry_increase': (
experiment_analysis['retry_count'] - baseline_analysis['retry_count']
),
'new_affected_operations': (
experiment_analysis['affected_operations']
- baseline_analysis['affected_operations']
),
'timeout_increase': (
experiment_analysis['timeout_count'] - baseline_analysis['timeout_count']
)
}
def _fetch_traces(self, service: str, start: datetime, end: datetime) -> List:
# Query Jaeger API for traces
params = {
'service': service,
'start': int(start.timestamp() * 1000000),
'end': int(end.timestamp() * 1000000),
'limit': 10000
}
response = requests.get(f"{self.jaeger_endpoint}/api/traces", params=params)
return response.json()['data']
def _has_error_span(self, trace: Dict) -> bool:
for span in trace['spans']:
if any(tag['key'] == 'error' and tag['value'] for tag in span.get('tags', [])):
return True
return False
def _count_retry_spans(self, trace: Dict) -> int:
retry_tags = ['retry', 'attempt', 'retry_count']
count = 0
for span in trace['spans']:
for tag in span.get('tags', []):
if tag['key'] in retry_tags:
count += 1
return count
def _has_timeout(self, trace: Dict) -> bool:
for span in trace['spans']:
for log in span.get('logs', []):
for field in log.get('fields', []):
if 'timeout' in field.get('value', '').lower():
return True
return False
def _circuit_breaker_triggered(self, trace: Dict) -> bool:
for span in trace['spans']:
for tag in span.get('tags', []):
if tag['key'] == 'circuit_breaker_state' and tag['value'] == 'open':
return True
return False
Log Aggregation and Analysis
# Loki query for chaos experiment logs
# Find errors during experiment window
{namespace="production", app="order-service"}
|= "error"
| json
| level="ERROR"
| __timestamp__ >= <experiment_start>
| __timestamp__ <= <experiment_end>
# Count errors by type
sum by (error_type) (
count_over_time(
{namespace="production", app="order-service"}
| json
| level="ERROR"
[5m]
)
)
# Detect timeout errors
{namespace="production"}
|~ "timeout|timed out|deadline exceeded"
| json
| line_format "{{.timestamp}} {{.service}} {{.message}}"
# Find circuit breaker events
{namespace="production"}
|= "circuit_breaker"
| json
| state="open"
---
# Python log analysis during chaos
import re
from datetime import datetime, timedelta
from typing import List, Dict
import requests
class ChaosLogAnalyser:
def __init__(self, loki_url: str):
self.loki_url = loki_url
def analyse_logs_during_experiment(
self,
namespace: str,
app: str,
start_time: datetime,
end_time: datetime
) -> Dict:
"""Analyse logs during chaos experiment."""
# Query for errors
error_query = f'{{namespace="{namespace}", app="{app}"}} |= "error" | json | level="ERROR"'
errors = self._query_range(error_query, start_time, end_time)
# Query for warnings
warn_query = f'{{namespace="{namespace}", app="{app}"}} |= "warn" | json | level="WARN"'
warnings = self._query_range(warn_query, start_time, end_time)
# Categorise errors
error_categories = self._categorise_errors(errors)
# Find anomalies
anomalies = self._detect_anomalies(errors, warnings)
return {
'total_errors': len(errors),
'total_warnings': len(warnings),
'error_categories': error_categories,
'anomalies': anomalies,
'timeline': self._create_timeline(errors, warnings)
}
def _categorise_errors(self, errors: List[Dict]) -> Dict:
"""Categorise errors by type."""
categories = {
'timeout': 0,
'connection_refused': 0,
'circuit_breaker': 0,
'resource_exhaustion': 0,
'other': 0
}
patterns = {
'timeout': r'timeout|timed out|deadline exceeded',
'connection_refused': r'connection refused|connection reset',
'circuit_breaker': r'circuit breaker|circuit open',
'resource_exhaustion': r'out of memory|too many open files|pool exhausted'
}
for error in errors:
message = error.get('message', '').lower()
categorised = False
for category, pattern in patterns.items():
if re.search(pattern, message):
categories[category] += 1
categorised = True
break
if not categorised:
categories['other'] += 1
return categories
def _detect_anomalies(self, errors: List, warnings: List) -> List[Dict]:
"""Detect anomalous log patterns."""
anomalies = []
# Spike detection: sudden increase in errors
time_windows = self._group_by_time_window(errors, window_minutes=1)
baseline_rate = sum(len(w) for w in time_windows[:5]) / 5 # First 5 minutes
for i, window in enumerate(time_windows[5:], start=5):
if len(window) > baseline_rate * 3: # 3x increase
anomalies.append({
'type': 'error_spike',
'time': i,
'count': len(window),
'baseline': baseline_rate
})
return anomalies
def _query_range(self, query: str, start: datetime, end: datetime) -> List:
"""Execute Loki range query."""
params = {
'query': query,
'start': int(start.timestamp() * 1000000000),
'end': int(end.timestamp() * 1000000000),
'limit': 5000
}
response = requests.get(f"{self.loki_url}/loki/api/v1/query_range", params=params)
results = response.json()['data']['result']
logs = []
for stream in results:
for value in stream['values']:
timestamp, message = value
logs.append({
'timestamp': datetime.fromtimestamp(int(timestamp) / 1000000000),
'message': message,
'labels': stream['stream']
})
return logs
def _group_by_time_window(self, logs: List, window_minutes: int) -> List[List]:
"""Group logs into time windows."""
if not logs:
return []
windows = []
current_window = []
window_start = logs[0]['timestamp']
window_duration = timedelta(minutes=window_minutes)
for log in logs:
if log['timestamp'] - window_start > window_duration:
windows.append(current_window)
current_window = [log]
window_start = log['timestamp']
else:
current_window.append(log)
if current_window:
windows.append(current_window)
return windows
def _create_timeline(self, errors: List, warnings: List) -> List[Dict]:
"""Create timeline of significant events."""
events = []
for error in errors:
events.append({
'timestamp': error['timestamp'],
'type': 'error',
'message': error['message']
})
for warning in warnings:
events.append({
'timestamp': warning['timestamp'],
'type': 'warning',
'message': warning['message']
})
# Sort by timestamp
events.sort(key=lambda x: x['timestamp'])
return events
Experiment Report Generation
# Generate comprehensive chaos experiment report
from dataclasses import dataclass
from typing import List, Dict
import json
@dataclass
class ExperimentReport:
experiment_name: str
hypothesis: str
start_time: datetime
end_time: datetime
blast_radius: Dict
# Metrics
baseline_metrics: Dict
experiment_metrics: Dict
recovery_metrics: Dict
# Observability data
trace_analysis: Dict
log_analysis: Dict
# Results
hypothesis_validated: bool
findings: List[str]
action_items: List[str]
def generate_report(self) -> str:
"""Generate markdown report."""
duration = (self.end_time - self.start_time).total_seconds() / 60
report = f"""# Chaos Experiment Report: {self.experiment_name}
## Experiment Details
**Date:** {self.start_time.strftime('%Y-%m-%d %H:%M UTC')}
**Duration:** {duration:.1f} minutes
**Blast Radius:** {json.dumps(self.blast_radius, indent=2)}
## Hypothesis
{self.hypothesis}
**Result:** {'✅ VALIDATED' if self.hypothesis_validated else '❌ INVALIDATED'}
## Metrics Comparison
### Request Rate
- **Baseline:** {self.baseline_metrics.get('request_rate', 'N/A')} req/s
- **During Experiment:** {self.experiment_metrics.get('request_rate', 'N/A')} req/s
- **After Recovery:** {self.recovery_metrics.get('request_rate', 'N/A')} req/s
### Error Rate
- **Baseline:** {self.baseline_metrics.get('error_rate', 0):.2%}
- **During Experiment:** {self.experiment_metrics.get('error_rate', 0):.2%}
- **After Recovery:** {self.recovery_metrics.get('error_rate', 0):.2%}
### Latency (P99)
- **Baseline:** {self.baseline_metrics.get('p99_latency', 'N/A')}ms
- **During Experiment:** {self.experiment_metrics.get('p99_latency', 'N/A')}ms
- **After Recovery:** {self.recovery_metrics.get('p99_latency', 'N/A')}ms
## Observability Analysis
### Distributed Tracing
- **Total Traces:** {self.trace_analysis.get('total_traces', 0)}
- **Error Count:** {self.trace_analysis.get('error_count', 0)}
- **Retry Events:** {self.trace_analysis.get('retry_count', 0)}
- **Circuit Breaker Activations:** {self.trace_analysis.get('circuit_breaker_open', 0)}
### Log Analysis
- **Total Errors:** {self.log_analysis.get('total_errors', 0)}
- **Total Warnings:** {self.log_analysis.get('total_warnings', 0)}
- **Error Categories:** {json.dumps(self.log_analysis.get('error_categories', {}), indent=2)}
## Key Findings
"""
for i, finding in enumerate(self.findings, 1):
report += f"{i}. {finding}\n"
report += "\n## Action Items\n\n"
for i, action in enumerate(self.action_items, 1):
report += f"- [ ] {action}\n"
return report
def save_to_file(self, filename: str):
"""Save report to markdown file."""
with open(filename, 'w') as f:
f.write(self.generate_report())
# Usage example
report = ExperimentReport(
experiment_name="Payment Service Pod Failure",
hypothesis="When 30% of payment-service pods fail, Istio retry logic maintains > 99.5% success rate",
start_time=datetime(2024, 1, 15, 10, 0),
end_time=datetime(2024, 1, 15, 10, 15),
blast_radius={
"environment": "production",
"namespace": "payments",
"affected_pods": "30%",
"duration": "5 minutes"
},
baseline_metrics={
"request_rate": 500,
"error_rate": 0.001,
"p99_latency": 180
},
experiment_metrics={
"request_rate": 485,
"error_rate": 0.003,
"p99_latency": 220
},
recovery_metrics={
"request_rate": 498,
"error_rate": 0.001,
"p99_latency": 185
},
trace_analysis={
"total_traces": 15000,
"error_count": 45,
"retry_count": 120,
"circuit_breaker_open": 0
},
log_analysis={
"total_errors": 42,
"total_warnings": 156,
"error_categories": {
"timeout": 5,
"connection_refused": 37,
"other": 0
}
},
hypothesis_validated=True,
findings=[
"Istio retry logic successfully handled pod failures with minimal user impact",
"Error rate remained below 0.5% throughout experiment",
"P99 latency increased by 22% but stayed within acceptable bounds",
"System fully recovered within 2 minutes of fault injection ending"
],
action_items=[
"Tune Istio retry timeout from 3s to 2s to reduce latency impact",
"Add connection pool monitoring to detect exhaustion earlier",
"Automate this experiment to run weekly in production"
]
)
report.save_to_file("chaos-experiment-2024-01-15.md")
Quick Reference
Chaos Experiment Workflow
| Phase | Duration | Actions | Validation |
|---|---|---|---|
| Planning | 1-2 days | Define hypothesis, select blast radius, identify metrics | Peer review, stakeholder approval |
| Preparation | 30 min | Set up monitoring, configure safeguards, run preflight checks | Health checks passing, alerts configured |
| Baseline | 15 min | Collect steady-state metrics before experiment | Metrics within normal range |
| Execution | 5-30 min | Inject fault, monitor continuously | Safeguards active, abort ready |
| Recovery | 15-30 min | Stop fault, verify system returns to steady state | All metrics back to baseline |
| Analysis | 1-2 hours | Compare baseline vs experiment, validate hypothesis | Report generated |
| Follow-up | 1 week | Implement improvements, document learnings | Action items tracked |
Common Experiment Types
| Experiment | Blast Radius | Duration | Key Metrics |
|---|---|---|---|
| Pod Failure | 10-30% pods | 2-5 min | Error rate, replica recovery time |
| Network Latency | 25-50% traffic | 5-15 min | P99 latency, timeout rate |
| CPU Stress | 20-40% pods | 10-20 min | Response time, autoscaling trigger |
| Memory Pressure | 10-20% pods | 5-10 min | OOM kills, pod restart rate |
| Zone Failure | 1 AZ (of 3+) | 10-30 min | Cross-AZ traffic, availability |
| Dependency Failure | 100% of dependency | 3-10 min | Circuit breaker activation, cache hit rate |
Tool Selection Guide
| Requirement | Recommended Tool | Reason |
|---|---|---|
| Kubernetes-only | Chaos Mesh | Native K8s integration, comprehensive |
| GitOps workflow | Litmus | Excellent CI/CD integration |
| Compliance/audit | Gremlin | Enterprise features, detailed reporting |
| AWS infrastructure | AWS FIS | Native AWS service integration |
| Multi-cloud | Gremlin or Litmus | Platform-agnostic |
| Cost-sensitive | Chaos Mesh or Litmus | Open source, no licensing |
Safety Checklist
- [ ] Hypothesis clearly defined
- [ ] Blast radius limited and documented
- [ ] Preflight health checks configured
- [ ] Abort conditions specified
- [ ] Manual abort procedure documented
- [ ] Monitoring and alerting active
- [ ] On-call engineer aware and available
- [ ] Stakeholders notified
- [ ] Rollback plan prepared
- [ ] Postflight validation defined
Common Issues and Solutions
Issue: Experiment Causes Cascading Failures
Symptoms:
- Blast radius expands beyond intended scope
- Multiple services affected
- Error rate exceeds abort thresholds
- Recovery takes longer than expected
Solutions:
- Implement circuit breakers on all service dependencies
# Istio circuit breaker trafficPolicy: outlierDetection: consecutiveErrors: 5 interval: 30s baseEjectionTime: 30s - Start with smaller blast radius (5-10% instead of 30%)
- Add dependency health checks to abort conditions
- Use bulkheads to isolate failure domains
- Increase resource limits temporarily during experiments
Issue: Unable to Validate Hypothesis
Symptoms:
- Results inconclusive
- Metrics don't show expected behaviour
- Missing observability data
- Hypothesis too vague
Solutions:
- Define specific, measurable hypothesis
- Bad: "Service should be resilient"
- Good: "Error rate < 1% when 30% pods fail for 5 minutes"
- Add instrumentation before experimenting
- Ensure traces include trace_id in logs
- Add custom metrics for business logic
- Configure log sampling appropriately
- Extend baseline collection period (30 min instead of 15 min)
- Run multiple iterations to account for variability
Issue: Production Experiments Too Risky
Symptoms:
- Stakeholders uncomfortable with production chaos
- Previous experiments caused incidents
- Lack of confidence in abort mechanisms
- Insufficient observability
Solutions:
- Progress through environments gradually
- Local development
- Integration environment
- Staging with production-like load
- Production with minimal blast radius
- Expand production scope gradually
- Implement feature flags for experiment control
if feature_flag.is_enabled('chaos_experiments', user_id): # Apply chaos only to flagged traffic - Use canary releases to limit experiment scope
- Run experiments during low-traffic periods initially
- Require explicit approval for production experiments
- Test abort mechanisms thoroughly in staging first
Issue: Experiments Don't Find Weaknesses
Symptoms:
- All experiments pass
- No improvements identified
- Team confidence may be false
- Not testing realistic scenarios
Solutions:
- Increase experiment severity
- Higher blast radius (30% → 50%)
- Longer duration (5 min → 15 min)
- Multiple simultaneous failures
- Test more realistic scenarios
# Combined network + resource stress apiVersion: chaos-mesh.org/v1alpha1 kind: Workflow metadata: name: realistic-failure spec: entry: combined-chaos templates: - name: combined-chaos templateType: Parallel children: - network-latency - cpu-stress - pod-failure - Challenge assumptions in hypothesis
- "Service handles 30% pod failure" → "Service handles 70% pod failure"
- Test rare but impactful scenarios
- Database connection pool exhaustion
- DNS resolution failures
- Dependency cascade failures
- Review incident history for real-world failure modes
Issue: Alert Noise During Experiments
Symptoms:
- Too many alerts fire
- On-call paged unnecessarily
- Hard to distinguish experiment from real incidents
- Alert fatigue
Solutions:
- Tag experiment resources
metadata: labels: chaos.enabled: "true" chaos.experiment: "pod-failure-2024-01-15" - Suppress alerts during experiments
# Alertmanager inhibition inhibit_rules: - source_match: chaos_experiment: "active" target_match: severity: "warning" equal: ['namespace', 'service'] - Create separate chaos alert channel
route: routes: - match: chaos: "true" receiver: chaos-team continue: true # Still send to normal channels - Communicate experiment schedule to team
- Use experiment annotations in Grafana to mark timeframes
Issue: Game Days Are Chaotic and Unproductive
Symptoms:
- Game day sessions lack structure
- No clear objectives or success criteria
- Team overwhelmed by multiple failures
- Learnings not documented
Solutions:
- Create structured game day runbook
# Game Day Runbook ## Objectives 1. Test incident response procedures 2. Validate RTO/RPO for database failure 3. Train new team members on escalation ## Scenario Timeline 09:00 - Intro and scenario briefing 09:15 - Inject database primary failure 09:20 - Teams respond, IC coordinates 09:45 - Validate recovery completed 10:00 - Debrief and retrospective ## Success Criteria - Database failover completes in < 5 minutes - No data loss - All team members know their roles - Runbook accurate and helpful - Assign roles clearly (Incident Commander, responders, observers)
- Start with one scenario before adding complexity
- Have facilitator who isn't responding to incident
- Schedule retrospective immediately after (while fresh)
- Document action items and assign owners
Issue: Chaos Experiments Not Integrated into CI/CD
Symptoms:
- Experiments run ad-hoc
- No automated regression testing for resilience
- Regressions introduced by new deployments
- Manual effort to run experiments
Solutions:
- Add chaos tests to CI pipeline
# GitLab CI example chaos-test: stage: test script: - kubectl apply -f chaos/network-latency.yaml - ./scripts/wait-for-experiment.sh - ./scripts/validate-metrics.sh - kubectl delete -f chaos/network-latency.yaml only: - merge_requests environment: name: staging - Use Litmus with GitOps
# Argo Workflow integration apiVersion: argoproj.io/v1alpha1 kind: Workflow metadata: name: chaos-pipeline spec: entrypoint: run-chaos-tests templates: - name: run-chaos-tests steps: - - name: deploy-app template: deploy - - name: pod-failure-test template: litmus-experiment arguments: parameters: - name: experiment value: "pod-delete" - - name: validate-results template: check-metrics - Schedule recurring experiments
# CronJob for weekly chaos apiVersion: batch/v1 kind: CronJob metadata: name: weekly-pod-chaos spec: schedule: "0 2 * * MON" # Every Monday 2am jobTemplate: spec: template: spec: containers: - name: chaos-runner image: chaostoolkit/chaostoolkit command: - chaos - run - /experiments/pod-failure.yaml - Fail builds if resilience tests fail
Related Topics
The following topics would complement this Chaos Engineering Practices cheatsheet:
-
Observability Patterns - Comprehensive guide to metrics, logs, and traces for validating chaos experiments and measuring system health
-
SLOs, SLIs, and Error Budgets - Detailed coverage of defining service level objectives and using error budgets to guide chaos engineering priorities
-
Kubernetes Advanced - Deep dive into Kubernetes resilience features like pod disruption budgets, health checks, and graceful shutdown
-
Istio/Service Mesh - Service mesh patterns for traffic management, circuit breaking, retries, and fault injection at the network level
-
Incident Management and Postmortems - Best practices for responding to real incidents and conducting blameless postmortems to drive improvements
-
Site Reliability Engineering (SRE) Practices - Broader SRE principles including toil reduction, capacity planning, and production readiness reviews that complement chaos engineering