Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Chaos Engineering Practices

A comprehensive guide to implementing chaos engineering for building resilient systems through controlled failure injection and experimentation.


Overview

Chaos Engineering is the discipline of experimenting on a system to build confidence in its capability to withstand turbulent conditions in production. Rather than waiting for failures to occur naturally, teams proactively inject faults to discover weaknesses before they impact users.

Continuous ValidationChaos Engineering CycleYesNoYesNoDefine Steady StateFormulate HypothesisDesign ExperimentDefine Blast RadiusSet SafeguardsRun ExperimentHypothesisValidated?Expand ScopeFix WeaknessDocument LearningsMonitorObservabilityHealth CheckFailing?Abort & RollbackContinuous ValidationChaos Engineering CycleYesNoYesNoDefine Steady StateFormulate HypothesisDesign ExperimentDefine Blast RadiusSet SafeguardsRun ExperimentHypothesisValidated?Expand ScopeFix WeaknessDocument LearningsMonitorObservabilityHealth CheckFailing?Abort & Rollback

Core Principles:

  1. Hypothesise about steady state - Define normal system behaviour
  2. Vary real-world events - Inject realistic failure scenarios
  3. Run experiments in production - Test where it matters most
  4. Automate experiments - Continuous chaos for continuous verification
  5. Minimise blast radius - Start small, expand gradually

Defining Steady State and Hypotheses

Key Concepts

Steady State represents normal system behaviour measured through business and technical metrics. It should focus on user-observable outcomes rather than internal system attributes.

Hypothesis Format:

Given [normal system state]
When [specific failure is injected]
Then [expected system behaviour]
And [user impact remains within acceptable bounds]

Steady State Metrics

Metric Type Examples Purpose
Business Metrics Orders/second, conversion rate, revenue Measure actual user impact
Application Metrics Request latency (p50/p99), error rate, throughput Track service health
Infrastructure Metrics CPU/memory usage, network bandwidth, disk I/O Identify resource constraints
User Experience Page load time, transaction success rate Direct impact measurement

Example Hypotheses

# Hypothesis 1: Service mesh retry resilience
steady_state:
  metric: "http_request_success_rate"
  threshold: "> 99.5%"
  duration: "5m"

hypothesis: |
  When 30% of payment-service pods fail,
  Then Istio retry logic will maintain > 99.5% success rate,
  And p99 latency will increase by < 50ms,
  And no user transactions will be lost.

validation_metrics:
  - name: "success_rate"
    query: "sum(rate(http_requests_total{status!~'5..'}[5m])) / sum(rate(http_requests_total[5m]))"
    threshold: "> 0.995"
  - name: "p99_latency"
    query: "histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))"
    threshold: "< 0.250"  # 200ms baseline + 50ms tolerance

# Hypothesis 2: Database connection pool exhaustion
steady_state:
  metric: "database_query_success_rate"
  threshold: "> 99.9%"
  duration: "10m"

hypothesis: |
  When database connection pool is exhausted (100 connections),
  Then application gracefully degrades with circuit breaker,
  And users see cached data rather than errors,
  And system auto-recovers when connections available.

validation_metrics:
  - name: "circuit_breaker_state"
    query: "circuit_breaker_open{service='api'}"
    threshold: "== 1"  # Should open
  - name: "cache_hit_rate"
    query: "rate(cache_hits[5m]) / rate(cache_requests[5m])"
    threshold: "> 0.80"

Steady State Definition Template

## Steady State Definition: [Service Name]

### Business Metrics
- **Orders Processed:** 1,200/hour (±10%)
- **Conversion Rate:** 3.5% (±0.5%)
- **Average Order Value:** £45 (±£5)

### Technical Metrics
- **Request Rate:** 500 req/s (±50 req/s)
- **Error Rate:** < 0.1% (5xx responses)
- **P50 Latency:** < 100ms
- **P99 Latency:** < 300ms
- **CPU Utilisation:** 40-60%
- **Memory Usage:** 2-4 GB

### Dependencies
- **Database:** Read latency < 10ms, Write latency < 25ms
- **Cache:** Hit rate > 85%
- **External API:** Response time < 500ms, Availability > 99.9%

### Acceptable Degradation Bounds
- Latency increase: < 50% from baseline
- Throughput decrease: < 20% from baseline
- Error rate: Must remain < 1%

Blast Radius and Scope Control

Key Concepts

Blast Radius defines the scope and potential impact of a chaos experiment. Starting with minimal blast radius and gradually expanding is critical for safe chaos engineering in production.

Blast Radius ProgressionSingle PodDev EnvironmentPod GroupStagingSingle AZStagingSingle PodProduction10% TrafficProductionSingle AZProductionRegionProductionBlast Radius ProgressionSingle PodDev EnvironmentPod GroupStagingSingle AZStagingSingle PodProduction10% TrafficProductionSingle AZProductionRegionProduction

Blast Radius Dimensions

Dimension Controls Examples
Environment Where experiment runs Dev → Staging → Production
Infrastructure Resources affected Single pod → Node → AZ → Region
Traffic Percentage of requests 1% → 10% → 50% → 100%
Time Duration of experiment 30s → 5m → 30m → Continuous
User Scope User population Internal → Beta users → All users

Blast Radius Configuration Examples

# Chaos Mesh - Pod failure with controlled scope
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: payment-service-failure
  namespace: chaos-testing
spec:
  action: pod-failure
  mode: fixed-percent  # Control blast radius
  value: "30"          # Affect 30% of pods
  duration: "2m"

  selector:
    namespaces:
      - production
    labelSelectors:
      app: payment-service
      version: v2.1

    # Additional filters to limit scope
    expressionSelectors:
      - key: chaos.enabled
        operator: In
        values: ["true"]

  # Schedule for controlled timing
  scheduler:
    cron: "@every 6h"

---
# Litmus - Network latency with gradual ramp-up
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: network-latency-rampup
  namespace: production
spec:
  engineState: active
  appinfo:
    appns: production
    applabel: app=checkout-service
    appkind: deployment

  chaosServiceAccount: litmus-admin
  experiments:
    - name: pod-network-latency
      spec:
        components:
          env:
            - name: NETWORK_LATENCY
              value: "2000"  # 2s latency
            - name: JITTER
              value: "500"   # ±500ms
            - name: TARGET_PODS
              value: "10%"   # Start with 10% blast radius
            - name: TOTAL_CHAOS_DURATION
              value: "300"   # 5 minutes
            - name: RAMP_TIME
              value: "60"    # Gradual ramp-up over 1 minute

Blast Radius Decision Matrix

Experiment Type Starting Scope Production-Ready Scope Notes
Pod Failure 1 pod in staging 10-30% of pods Ensure replicas > 3
Network Latency 1% traffic 25-50% traffic Monitor p99 latency
CPU Stress Single pod, 50% load 30% pods, 80% load Watch for cascade
Memory Pressure 1 pod, 70% memory 20% pods, 85% memory Risk of OOM kills
DNS Failure Non-critical service Critical service, 10% pods High impact potential
Zone Failure Staging only Production with 3+ zones Major blast radius

Common Chaos Experiments

Pod and Container Failures

# Chaos Mesh - Pod Kill
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: pod-kill-experiment
spec:
  action: pod-kill
  mode: one
  selector:
    namespaces:
      - production
    labelSelectors:
      app: api-service
  scheduler:
    cron: "@every 4h"

---
# Chaos Mesh - Container Kill (partial pod failure)
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: container-kill-experiment
spec:
  action: container-kill
  mode: fixed
  value: "2"
  containerNames:
    - sidecar-proxy  # Test sidecar resilience
  selector:
    namespaces:
      - production
    labelSelectors:
      app: payment-service
  duration: "5m"

Network Chaos

# Chaos Mesh - Network Partition
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: network-partition
spec:
  action: partition
  mode: all
  selector:
    namespaces:
      - production
    labelSelectors:
      app: order-service
  direction: to
  target:
    selector:
      namespaces:
        - production
      labelSelectors:
        app: inventory-service
  duration: "3m"

---
# Chaos Mesh - Network Delay
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: network-latency-injection
spec:
  action: delay
  mode: fixed-percent
  value: "50"
  delay:
    latency: "250ms"
    correlation: "50"  # 50% correlation between packets
    jitter: "50ms"
  selector:
    namespaces:
      - production
    labelSelectors:
      app: checkout-service
  duration: "10m"

---
# Chaos Mesh - Packet Loss
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: packet-loss
spec:
  action: loss
  mode: one
  loss:
    loss: "25"        # 25% packet loss
    correlation: "25"
  selector:
    namespaces:
      - production
    labelSelectors:
      app: notification-service
  direction: to
  target:
    selector:
      namespaces:
        - production
      labelSelectors:
        app: email-gateway
  duration: "5m"

---
# Chaos Mesh - Bandwidth Limitation
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: bandwidth-limit
spec:
  action: bandwidth
  mode: all
  bandwidth:
    rate: "1mbps"      # Limit to 1 Mbps
    limit: 20000       # Buffer size
    buffer: 10000      # Peakrate buffer
  selector:
    namespaces:
      - production
    labelSelectors:
      app: video-streaming
  duration: "15m"

Resource Stress

# Chaos Mesh - CPU Stress
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
  name: cpu-stress
spec:
  mode: fixed-percent
  value: "40"
  stressors:
    cpu:
      workers: 4
      load: 80      # 80% CPU load
  selector:
    namespaces:
      - production
    labelSelectors:
      app: analytics-service
  duration: "10m"

---
# Chaos Mesh - Memory Stress
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
  name: memory-pressure
spec:
  mode: one
  stressors:
    memory:
      workers: 4
      size: "512MB"  # Allocate 512MB
  selector:
    namespaces:
      - production
    labelSelectors:
      app: cache-service
  duration: "5m"

---
# Litmus - Disk Fill
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: disk-fill
spec:
  engineState: active
  appinfo:
    appns: production
    applabel: app=database
    appkind: statefulset
  chaosServiceAccount: litmus-admin
  experiments:
    - name: disk-fill
      spec:
        components:
          env:
            - name: FILL_PERCENTAGE
              value: "80"       # Fill disk to 80%
            - name: TARGET_CONTAINER
              value: "postgres"
            - name: TOTAL_CHAOS_DURATION
              value: "300"

Application-Level Failures

# Chaos Mesh - HTTP Abort (Service unavailable)
apiVersion: chaos-mesh.org/v1alpha1
kind: HTTPChaos
metadata:
  name: http-abort
spec:
  mode: fixed-percent
  value: "25"
  target: Request
  port: 8080
  method: POST
  path: /api/v1/orders
  abort: true
  selector:
    namespaces:
      - production
    labelSelectors:
      app: order-service
  duration: "5m"

---
# Chaos Mesh - HTTP Delay
apiVersion: chaos-mesh.org/v1alpha1
kind: HTTPChaos
metadata:
  name: http-latency
spec:
  mode: all
  target: Request
  port: 8080
  path: /api/v1/payments
  delay: "3s"
  selector:
    namespaces:
      - production
    labelSelectors:
      app: payment-gateway
  duration: "10m"

---
# Gremlin - DNS Failure
gremlin attack create dns \
  --hostname database.internal.example.com \
  --length 300 \
  --tags "service:api,env:production" \
  --target-type container \
  --target-percent 20

Infrastructure Failures

# Chaos Mesh - Kernel I/O Chaos
apiVersion: chaos-mesh.org/v1alpha1
kind: IOChaos
metadata:
  name: io-latency
spec:
  action: latency
  mode: one
  volumePath: /var/lib/postgresql/data
  path: /var/lib/postgresql/data/**/*
  delay: "100ms"
  percent: 50
  selector:
    namespaces:
      - production
    labelSelectors:
      app: postgres
  duration: "10m"

---
# Litmus - Node Drain (simulates node failure)
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: node-drain-chaos
spec:
  engineState: active
  chaosServiceAccount: litmus-admin
  experiments:
    - name: node-drain
      spec:
        components:
          env:
            - name: TARGET_NODE
              value: "ip-10-0-1-50.eu-west-1.compute.internal"
            - name: TOTAL_CHAOS_DURATION
              value: "600"  # 10 minutes

---
# AWS Fault Injection Simulator - AZ outage
{
  "description": "Simulate AZ failure",
  "targets": {
    "Subnets": {
      "resourceType": "aws:ec2:subnet",
      "resourceArns": [
        "arn:aws:ec2:eu-west-1:123456789012:subnet/subnet-abc123"
      ],
      "selectionMode": "ALL"
    }
  },
  "actions": {
    "BlockSubnet": {
      "actionId": "aws:network:disrupt-connectivity",
      "parameters": {
        "duration": "PT5M",
        "scope": "all"
      },
      "targets": {
        "Subnets": "Subnets"
      }
    }
  },
  "stopConditions": [
    {
      "source": "aws:cloudwatch:alarm",
      "value": "arn:aws:cloudwatch:eu-west-1:123456789012:alarm:HighErrorRate"
    }
  ],
  "roleArn": "arn:aws:iam::123456789012:role/FISRole"
}

Chaos Engineering Tools

Chaos Mesh (Kubernetes-native)

# Install Chaos Mesh
helm repo add chaos-mesh https://charts.chaos-mesh.org
helm install chaos-mesh chaos-mesh/chaos-mesh \
  --namespace=chaos-mesh \
  --create-namespace \
  --set dashboard.create=true

# Access dashboard
kubectl port-forward -n chaos-mesh svc/chaos-dashboard 2333:2333

# Create experiment via CLI
kubectl apply -f - <<EOF
apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
  name: payment-resilience-test
  namespace: chaos-testing
spec:
  entry: entry
  templates:
    - name: entry
      templateType: Serial
      children:
        - network-delay
        - pod-failure
        - validate-recovery

    - name: network-delay
      templateType: NetworkChaos
      deadline: 5m
      networkChaos:
        action: delay
        mode: fixed-percent
        value: "30"
        delay:
          latency: "500ms"
        selector:
          namespaces: [production]
          labelSelectors:
            app: payment-service

    - name: pod-failure
      templateType: PodChaos
      deadline: 3m
      podChaos:
        action: pod-kill
        mode: fixed
        value: "2"
        selector:
          namespaces: [production]
          labelSelectors:
            app: payment-service

    - name: validate-recovery
      templateType: Suspend
      deadline: 10m
EOF

# List experiments
kubectl get podchaos,networkchaos,stresschaos -A

# Pause experiment
kubectl annotate podchaos pod-kill-experiment experiment.chaos-mesh.org/pause=true

# Resume experiment
kubectl annotate podchaos pod-kill-experiment experiment.chaos-mesh.org/pause-

# Delete experiment
kubectl delete podchaos pod-kill-experiment

Litmus Chaos

# Install Litmus
kubectl apply -f https://litmuschaos.github.io/litmus/litmus-operator-latest.yaml

# Install chaos experiments
kubectl apply -f https://hub.litmuschaos.io/api/chaos/master?file=charts/generic/experiments.yaml

# Create service account
kubectl apply -f - <<EOF
apiVersion: v1
kind: ServiceAccount
metadata:
  name: litmus-admin
  namespace: production
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: litmus-admin
rules:
  - apiGroups: [""]
    resources: ["pods", "events"]
    verbs: ["create", "delete", "get", "list", "patch", "update"]
  - apiGroups: ["apps"]
    resources: ["deployments", "statefulsets"]
    verbs: ["get", "list", "update"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: litmus-admin
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: litmus-admin
subjects:
  - kind: ServiceAccount
    name: litmus-admin
    namespace: production
EOF

# Run experiment
kubectl apply -f chaos-engine.yaml

# Monitor experiment
kubectl describe chaosengine network-latency-rampup -n production

# View results
kubectl logs -n production \
  $(kubectl get pods -n production -l chaosUID -o jsonpath='{.items[0].metadata.name}') \
  -c chaos-runner

# Abort experiment
kubectl patch chaosengine network-latency-rampup -n production \
  --type merge -p '{"spec":{"engineState":"stop"}}'

Gremlin (SaaS Platform)

# Install Gremlin agent
helm repo add gremlin https://helm.gremlin.com
helm install gremlin gremlin/gremlin \
  --namespace gremlin \
  --create-namespace \
  --set gremlin.teamID=$GREMLIN_TEAM_ID \
  --set gremlin.teamSecret=$GREMLIN_TEAM_SECRET \
  --set gremlin.clusterID=$CLUSTER_NAME

# Create CPU stress attack (via CLI)
gremlin attack create cpu \
  --cores 2 \
  --percent 80 \
  --length 300 \
  --tags "service:payment,env:production" \
  --target-type container \
  --target-percent 25

# Create network latency attack
gremlin attack create latency \
  --latency 500 \
  --length 600 \
  --tags "service:checkout" \
  --target-type container

# Create blackhole (network partition)
gremlin attack create blackhole \
  --hostname database.internal \
  --port 5432 \
  --length 180 \
  --tags "service:api"

# List active attacks
gremlin attack list --active

# Halt attack
gremlin attack halt <attack-id>

# Create scenario (multi-step experiment)
gremlin scenario create \
  --name "Database Failover Test" \
  --schedule "0 2 * * MON" \
  --steps '[
    {
      "type": "latency",
      "target": {"tags": ["role:primary-db"]},
      "args": {"latency": 1000, "length": 300}
    },
    {
      "type": "shutdown",
      "target": {"tags": ["role:primary-db"]},
      "args": {"delay": 10}
    }
  ]'

AWS Fault Injection Simulator (FIS)

# Create experiment template
aws fis create-experiment-template \
  --cli-input-json file://fis-template.json

# Start experiment
EXPERIMENT_ID=$(aws fis start-experiment \
  --experiment-template-id EXT123456 \
  --query 'experiment.id' \
  --output text)

# Monitor experiment
aws fis get-experiment --id $EXPERIMENT_ID

# Stop experiment
aws fis stop-experiment --id $EXPERIMENT_ID

# Example: EC2 instance termination
{
  "description": "Test auto-scaling response",
  "targets": {
    "Instances": {
      "resourceType": "aws:ec2:instance",
      "resourceTags": {
        "Environment": "production",
        "ChaosReady": "true"
      },
      "filters": [
        {
          "path": "State.Name",
          "values": ["running"]
        }
      ],
      "selectionMode": "COUNT(2)"
    }
  },
  "actions": {
    "TerminateInstances": {
      "actionId": "aws:ec2:terminate-instances",
      "targets": {
        "Instances": "Instances"
      }
    }
  },
  "stopConditions": [
    {
      "source": "aws:cloudwatch:alarm",
      "value": "arn:aws:cloudwatch:eu-west-1:123456789012:alarm:UnhealthyHosts"
    }
  ],
  "roleArn": "arn:aws:iam::123456789012:role/FISRole"
}

Tool Comparison

Feature Chaos Mesh Litmus Gremlin AWS FIS
Platform Kubernetes Kubernetes Multi-platform AWS only
Cost Free Free Commercial Pay-per-use
Deployment Self-hosted Self-hosted SaaS + Agent Managed
Experiment Types 10+ types 30+ experiments 15+ attacks AWS resources
GUI Dashboard ChaosCenter Web platform AWS Console
Workflows Native Native Scenarios Templates
Observability Prometheus metrics Built-in Integrated CloudWatch
RBAC Kubernetes RBAC Kubernetes RBAC Built-in IAM
Best For K8s chaos K8s + GitOps Enterprise, compliance AWS infrastructure

Safeguards and Abort Conditions

Key Concepts

Safeguards prevent chaos experiments from causing unacceptable harm. Every experiment must have automated abort conditions and manual stop mechanisms.

PassFailHealthyUnhealthyNoYesNoYesRecoveredNot RecoveredStart ExperimentPre-flightChecksInject FaultAbort -Prerequisites NotMetMonitorHealth ChecksAbortConditionTriggered?Abort - Health CheckFailedDurationComplete?Abort - ConditionMetStop Fault InjectionVerifyRecoveryCompleteSuccessfullyAlert - ManualInterventionCleanupPassFailHealthyUnhealthyNoYesNoYesRecoveredNot RecoveredStart ExperimentPre-flightChecksInject FaultAbort -Prerequisites NotMetMonitorHealth ChecksAbortConditionTriggered?Abort - Health CheckFailedDurationComplete?Abort - ConditionMetStop Fault InjectionVerifyRecoveryCompleteSuccessfullyAlert - ManualInterventionCleanup

Pre-flight Checks

# Chaos Mesh - StatusCheck (custom resource)
apiVersion: chaos-mesh.org/v1alpha1
kind: StatusCheck
metadata:
  name: preflight-checks
spec:
  mode: Synchronous
  type: HTTP
  intervalSeconds: 5
  timeoutSeconds: 10
  successThreshold: 3
  failureThreshold: 1

  http:
    url: http://health-check-service/readiness
    method: GET
    criteria:
      statusCode: "200"

  # Use in workflow
---
apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
  name: safe-experiment
spec:
  entry: check-then-chaos
  templates:
    - name: check-then-chaos
      templateType: Serial
      children:
        - preflight-validation
        - run-chaos
        - postflight-validation

    - name: preflight-validation
      templateType: StatusCheck
      statusCheck:
        mode: Synchronous
        type: HTTP
        http:
          url: http://api-service/health
          criteria:
            statusCode: "200"
        successThreshold: 3
        failureThreshold: 1

    - name: run-chaos
      templateType: PodChaos
      # ... chaos spec

Health Check Examples

# Python health check for experiment safety
import requests
import prometheus_client as prom
from typing import Dict, List, Tuple

class ChaosHealthCheck:
    def __init__(self, prometheus_url: str):
        self.prom_url = prometheus_url

    def check_error_rate(self, threshold: float = 0.01) -> Tuple[bool, str]:
        """Verify error rate is below threshold."""
        query = """
        sum(rate(http_requests_total{status=~"5.."}[5m]))
        /
        sum(rate(http_requests_total[5m]))
        """
        result = self._query_prometheus(query)

        if not result:
            return False, "Unable to query error rate"

        error_rate = result[0]['value'][1]
        if float(error_rate) > threshold:
            return False, f"Error rate {error_rate} exceeds threshold {threshold}"

        return True, "Error rate acceptable"

    def check_latency(self, p99_threshold_ms: int = 1000) -> Tuple[bool, str]:
        """Verify p99 latency is below threshold."""
        query = f"""
        histogram_quantile(0.99,
          sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
        ) * 1000
        """
        result = self._query_prometheus(query)

        if not result:
            return False, "Unable to query latency"

        p99_latency = float(result[0]['value'][1])
        if p99_latency > p99_threshold_ms:
            return False, f"P99 latency {p99_latency}ms exceeds {p99_threshold_ms}ms"

        return True, f"P99 latency {p99_latency}ms acceptable"

    def check_service_availability(self, min_replicas: int = 2) -> Tuple[bool, str]:
        """Verify minimum number of healthy replicas."""
        query = f"""
        sum(up{{job="api-service"}})
        """
        result = self._query_prometheus(query)

        if not result:
            return False, "Unable to query service availability"

        available = int(float(result[0]['value'][1]))
        if available < min_replicas:
            return False, f"Only {available} replicas available, need {min_replicas}"

        return True, f"{available} replicas available"

    def check_dependency_health(self, dependencies: List[str]) -> Tuple[bool, str]:
        """Verify all dependencies are healthy."""
        for dep in dependencies:
            query = f'up{{job="{dep}"}}'
            result = self._query_prometheus(query)

            if not result or float(result[0]['value'][1]) != 1:
                return False, f"Dependency {dep} is unhealthy"

        return True, "All dependencies healthy"

    def run_preflight_checks(self) -> bool:
        """Run all preflight checks before starting experiment."""
        checks = [
            self.check_error_rate(threshold=0.005),  # 0.5% max
            self.check_latency(p99_threshold_ms=500),
            self.check_service_availability(min_replicas=3),
            self.check_dependency_health(['database', 'cache', 'auth-service'])
        ]

        for passed, message in checks:
            print(f"{'✓' if passed else '✗'} {message}")
            if not passed:
                print("❌ Preflight checks failed. Aborting experiment.")
                return False

        print("✅ All preflight checks passed. Safe to proceed.")
        return True

    def _query_prometheus(self, query: str) -> list:
        """Execute Prometheus query."""
        try:
            response = requests.get(
                f"{self.prom_url}/api/v1/query",
                params={'query': query},
                timeout=10
            )
            return response.json()['data']['result']
        except Exception as e:
            print(f"Error querying Prometheus: {e}")
            return []

# Usage
health_check = ChaosHealthCheck('http://prometheus:9090')
if health_check.run_preflight_checks():
    # Start chaos experiment
    pass

Abort Conditions

# AWS FIS - CloudWatch Alarm Stop Condition
stopConditions:
  - source: aws:cloudwatch:alarm
    value: arn:aws:cloudwatch:eu-west-1:123456789012:alarm:HighErrorRate
  - source: aws:cloudwatch:alarm
    value: arn:aws:cloudwatch:eu-west-1:123456789012:alarm:HighLatency
  - source: aws:cloudwatch:alarm
    value: arn:aws:cloudwatch:eu-west-1:123456789012:alarm:LowAvailability

---
# Chaos Mesh - StatusCheck for abort
apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
  name: chaos-with-abort
spec:
  entry: monitored-chaos
  templates:
    - name: monitored-chaos
      templateType: Parallel
      children:
        - inject-fault
        - monitor-health

    - name: inject-fault
      templateType: PodChaos
      deadline: 10m
      abortWithStatusCheck: health-monitor  # References monitor
      podChaos:
        action: pod-kill
        mode: fixed-percent
        value: "20"
        selector:
          namespaces: [production]
          labelSelectors:
            app: order-service

    - name: monitor-health
      templateType: StatusCheck
      statusCheck:
        mode: Continuous
        type: HTTP
        intervalSeconds: 10
        timeoutSeconds: 5
        failureThreshold: 3  # Abort after 3 consecutive failures
        http:
          url: http://order-service/health
          method: GET
          criteria:
            statusCode: "200"

---
# Custom abort controller in Python
import time
import requests
from datetime import datetime, timedelta

class ChaosAbortController:
    def __init__(self, experiment_name: str, prometheus_url: str):
        self.experiment_name = experiment_name
        self.prom_url = prometheus_url
        self.start_time = datetime.now()
        self.max_duration = timedelta(minutes=15)

    def should_abort(self) -> Tuple[bool, str]:
        """Check if experiment should be aborted."""

        # Timeout check
        if datetime.now() - self.start_time > self.max_duration:
            return True, "Maximum duration exceeded"

        # Error rate check
        error_rate = self._get_error_rate()
        if error_rate and error_rate > 0.05:  # 5% threshold
            return True, f"Error rate {error_rate:.2%} exceeds 5%"

        # Latency check
        p99_latency = self._get_p99_latency()
        if p99_latency and p99_latency > 2000:  # 2s threshold
            return True, f"P99 latency {p99_latency}ms exceeds 2000ms"

        # Business metric check
        throughput = self._get_throughput()
        if throughput and throughput < 100:  # Minimum req/s
            return True, f"Throughput {throughput} req/s below minimum 100 req/s"

        # Error budget check
        error_budget_remaining = self._get_error_budget_remaining()
        if error_budget_remaining is not None and error_budget_remaining < 0.10:
            return True, f"Error budget at {error_budget_remaining:.1%}, below 10% threshold"

        return False, "All checks passed"

    def monitor_and_abort(self, check_interval: int = 30):
        """Continuously monitor and abort if needed."""
        while True:
            should_abort, reason = self.should_abort()

            if should_abort:
                print(f"🛑 ABORTING EXPERIMENT: {reason}")
                self._abort_experiment()
                return

            print(f"✓ Health check passed: {reason}")
            time.sleep(check_interval)

    def _abort_experiment(self):
        """Abort the chaos experiment."""
        # Example for Chaos Mesh
        import subprocess
        subprocess.run([
            'kubectl', 'delete', 'podchaos', self.experiment_name,
            '--namespace', 'chaos-testing'
        ])

    def _get_error_rate(self) -> float:
        query = """
        sum(rate(http_requests_total{status=~"5.."}[1m]))
        / sum(rate(http_requests_total[1m]))
        """
        return self._query_single_value(query)

    def _get_p99_latency(self) -> float:
        query = """
        histogram_quantile(0.99,
          sum(rate(http_request_duration_seconds_bucket[1m])) by (le)
        ) * 1000
        """
        return self._query_single_value(query)

    def _get_throughput(self) -> float:
        query = "sum(rate(http_requests_total[1m]))"
        return self._query_single_value(query)

    def _get_error_budget_remaining(self) -> float:
        query = """
        1 - (
          (1 - (
            sum(rate(http_requests_total{status!~"5.."}[30d]))
            / sum(rate(http_requests_total[30d]))
          ))
          / (1 - 0.999)
        )
        """
        return self._query_single_value(query)

    def _query_single_value(self, query: str) -> float:
        try:
            response = requests.get(
                f"{self.prom_url}/api/v1/query",
                params={'query': query},
                timeout=5
            )
            result = response.json()['data']['result']
            if result:
                return float(result[0]['value'][1])
        except:
            pass
        return None

Circuit Breaker Pattern

# Istio DestinationRule with circuit breaker
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: payment-service-circuit-breaker
spec:
  host: payment-service
  trafficPolicy:
    connectionPool:
      tcp:
        maxConnections: 100
      http:
        http1MaxPendingRequests: 50
        http2MaxRequests: 100
        maxRequestsPerConnection: 2

    outlierDetection:
      consecutiveErrors: 5
      interval: 30s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
      minHealthPercent: 40

    # Automatically eject unhealthy pods during chaos
    loadBalancer:
      simple: LEAST_REQUEST

Observability Validation During Experiments

Key Concepts

Observability is critical for validating chaos experiments. You must monitor system behaviour before, during, and after fault injection to validate hypotheses and detect anomalies.

AnalysisRecovery PhaseExperiment PhaseBaseline CollectionYesNoEstablish BaselineT-15min to T0Record NormalBehaviourStart FaultInjectionT0Monitor MetricsT0 to T+NCapture Events& LogsStop FaultT+NMonitor RecoveryT+N to T+N+30Verify Steady StateRestoredCompare Baselinevs ExperimentValidate HypothesisHypothesisConfirmed?Document SuccessInvestigate GapAnalysisRecovery PhaseExperiment PhaseBaseline CollectionYesNoEstablish BaselineT-15min to T0Record NormalBehaviourStart FaultInjectionT0Monitor MetricsT0 to T+NCapture Events& LogsStop FaultT+NMonitor RecoveryT+N to T+N+30Verify Steady StateRestoredCompare Baselinevs ExperimentValidate HypothesisHypothesisConfirmed?Document SuccessInvestigate Gap

Metrics to Monitor

# Prometheus recording rules for chaos experiments
groups:
  - name: chaos_experiment_metrics
    interval: 10s
    rules:
      # Request metrics
      - record: chaos:request_rate:sum
        expr: sum(rate(http_requests_total[1m])) by (service)

      - record: chaos:error_rate:ratio
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[1m])) by (service)
          /
          sum(rate(http_requests_total[1m])) by (service)

      # Latency metrics
      - record: chaos:latency:p50
        expr: |
          histogram_quantile(0.50,
            sum(rate(http_request_duration_seconds_bucket[1m])) by (service, le)
          )

      - record: chaos:latency:p95
        expr: |
          histogram_quantile(0.95,
            sum(rate(http_request_duration_seconds_bucket[1m])) by (service, le)
          )

      - record: chaos:latency:p99
        expr: |
          histogram_quantile(0.99,
            sum(rate(http_request_duration_seconds_bucket[1m])) by (service, le)
          )

      # Resource metrics
      - record: chaos:cpu_usage:avg
        expr: avg(rate(container_cpu_usage_seconds_total[1m])) by (pod, container)

      - record: chaos:memory_usage:bytes
        expr: avg(container_memory_working_set_bytes) by (pod, container)

      # Infrastructure metrics
      - record: chaos:pod_ready:count
        expr: sum(kube_pod_status_ready{condition="true"}) by (namespace, deployment)

      - record: chaos:pod_restarts:rate
        expr: sum(rate(kube_pod_container_status_restarts_total[5m])) by (namespace, pod)

---
# Grafana dashboard for experiment monitoring
{
  "dashboard": {
    "title": "Chaos Experiment Dashboard",
    "panels": [
      {
        "title": "Request Rate (Baseline vs Experiment)",
        "targets": [
          {
            "expr": "chaos:request_rate:sum{service='$service'}",
            "legendFormat": "Current"
          },
          {
            "expr": "avg_over_time(chaos:request_rate:sum{service='$service'}[15m] offset 30m)",
            "legendFormat": "Baseline"
          }
        ]
      },
      {
        "title": "Error Rate",
        "targets": [
          {
            "expr": "chaos:error_rate:ratio{service='$service'} * 100",
            "legendFormat": "Error %"
          }
        ],
        "alert": {
          "conditions": [
            {
              "evaluator": {
                "params": [5],
                "type": "gt"
              },
              "query": {"params": ["A", "5m", "now"]}
            }
          ]
        }
      },
      {
        "title": "Latency Percentiles",
        "targets": [
          {
            "expr": "chaos:latency:p50{service='$service'} * 1000",
            "legendFormat": "P50"
          },
          {
            "expr": "chaos:latency:p95{service='$service'} * 1000",
            "legendFormat": "P95"
          },
          {
            "expr": "chaos:latency:p99{service='$service'} * 1000",
            "legendFormat": "P99"
          }
        ]
      }
    ],
    "annotations": {
      "list": [
        {
          "name": "Chaos Events",
          "datasource": "Prometheus",
          "expr": "ALERTS{alertname=~'Chaos.*'}",
          "iconColor": "red"
        },
        {
          "name": "Deployments",
          "datasource": "Prometheus",
          "expr": "changes(kube_deployment_status_observed_generation[30s]) > 0",
          "iconColor": "blue"
        }
      ]
    }
  }
}

Distributed Tracing Validation

# Analyse trace data during chaos experiments
from jaeger_client import Config
from opentelemetry import trace
import time

class ChaosTraceAnalyser:
    def __init__(self, jaeger_endpoint: str):
        self.jaeger_endpoint = jaeger_endpoint

    def analyse_traces_during_experiment(
        self,
        service_name: str,
        start_time: datetime,
        end_time: datetime
    ) -> Dict:
        """Analyse trace patterns during chaos experiment."""

        traces = self._fetch_traces(service_name, start_time, end_time)

        analysis = {
            'total_traces': len(traces),
            'error_count': 0,
            'retry_count': 0,
            'timeout_count': 0,
            'circuit_breaker_open': 0,
            'latency_spikes': [],
            'affected_operations': set()
        }

        for trace in traces:
            # Check for errors
            if self._has_error_span(trace):
                analysis['error_count'] += 1
                analysis['affected_operations'].add(trace['operationName'])

            # Detect retries
            retry_count = self._count_retry_spans(trace)
            if retry_count > 0:
                analysis['retry_count'] += retry_count

            # Detect timeouts
            if self._has_timeout(trace):
                analysis['timeout_count'] += 1

            # Detect circuit breaker
            if self._circuit_breaker_triggered(trace):
                analysis['circuit_breaker_open'] += 1

            # Identify latency spikes
            duration_ms = trace['duration'] / 1000
            if duration_ms > 1000:  # > 1s
                analysis['latency_spikes'].append({
                    'trace_id': trace['traceID'],
                    'duration_ms': duration_ms,
                    'operation': trace['operationName']
                })

        return analysis

    def compare_baseline_vs_experiment(
        self,
        baseline_analysis: Dict,
        experiment_analysis: Dict
    ) -> Dict:
        """Compare trace patterns between baseline and experiment."""

        return {
            'error_rate_delta': (
                experiment_analysis['error_count'] / experiment_analysis['total_traces']
                - baseline_analysis['error_count'] / baseline_analysis['total_traces']
            ),
            'retry_increase': (
                experiment_analysis['retry_count'] - baseline_analysis['retry_count']
            ),
            'new_affected_operations': (
                experiment_analysis['affected_operations']
                - baseline_analysis['affected_operations']
            ),
            'timeout_increase': (
                experiment_analysis['timeout_count'] - baseline_analysis['timeout_count']
            )
        }

    def _fetch_traces(self, service: str, start: datetime, end: datetime) -> List:
        # Query Jaeger API for traces
        params = {
            'service': service,
            'start': int(start.timestamp() * 1000000),
            'end': int(end.timestamp() * 1000000),
            'limit': 10000
        }
        response = requests.get(f"{self.jaeger_endpoint}/api/traces", params=params)
        return response.json()['data']

    def _has_error_span(self, trace: Dict) -> bool:
        for span in trace['spans']:
            if any(tag['key'] == 'error' and tag['value'] for tag in span.get('tags', [])):
                return True
        return False

    def _count_retry_spans(self, trace: Dict) -> int:
        retry_tags = ['retry', 'attempt', 'retry_count']
        count = 0
        for span in trace['spans']:
            for tag in span.get('tags', []):
                if tag['key'] in retry_tags:
                    count += 1
        return count

    def _has_timeout(self, trace: Dict) -> bool:
        for span in trace['spans']:
            for log in span.get('logs', []):
                for field in log.get('fields', []):
                    if 'timeout' in field.get('value', '').lower():
                        return True
        return False

    def _circuit_breaker_triggered(self, trace: Dict) -> bool:
        for span in trace['spans']:
            for tag in span.get('tags', []):
                if tag['key'] == 'circuit_breaker_state' and tag['value'] == 'open':
                    return True
        return False

Log Aggregation and Analysis

# Loki query for chaos experiment logs
# Find errors during experiment window
{namespace="production", app="order-service"}
  |= "error"
  | json
  | level="ERROR"
  | __timestamp__ >= <experiment_start>
  | __timestamp__ <= <experiment_end>

# Count errors by type
sum by (error_type) (
  count_over_time(
    {namespace="production", app="order-service"}
      | json
      | level="ERROR"
      [5m]
  )
)

# Detect timeout errors
{namespace="production"}
  |~ "timeout|timed out|deadline exceeded"
  | json
  | line_format "{{.timestamp}} {{.service}} {{.message}}"

# Find circuit breaker events
{namespace="production"}
  |= "circuit_breaker"
  | json
  | state="open"

---
# Python log analysis during chaos
import re
from datetime import datetime, timedelta
from typing import List, Dict
import requests

class ChaosLogAnalyser:
    def __init__(self, loki_url: str):
        self.loki_url = loki_url

    def analyse_logs_during_experiment(
        self,
        namespace: str,
        app: str,
        start_time: datetime,
        end_time: datetime
    ) -> Dict:
        """Analyse logs during chaos experiment."""

        # Query for errors
        error_query = f'{{namespace="{namespace}", app="{app}"}} |= "error" | json | level="ERROR"'
        errors = self._query_range(error_query, start_time, end_time)

        # Query for warnings
        warn_query = f'{{namespace="{namespace}", app="{app}"}} |= "warn" | json | level="WARN"'
        warnings = self._query_range(warn_query, start_time, end_time)

        # Categorise errors
        error_categories = self._categorise_errors(errors)

        # Find anomalies
        anomalies = self._detect_anomalies(errors, warnings)

        return {
            'total_errors': len(errors),
            'total_warnings': len(warnings),
            'error_categories': error_categories,
            'anomalies': anomalies,
            'timeline': self._create_timeline(errors, warnings)
        }

    def _categorise_errors(self, errors: List[Dict]) -> Dict:
        """Categorise errors by type."""
        categories = {
            'timeout': 0,
            'connection_refused': 0,
            'circuit_breaker': 0,
            'resource_exhaustion': 0,
            'other': 0
        }

        patterns = {
            'timeout': r'timeout|timed out|deadline exceeded',
            'connection_refused': r'connection refused|connection reset',
            'circuit_breaker': r'circuit breaker|circuit open',
            'resource_exhaustion': r'out of memory|too many open files|pool exhausted'
        }

        for error in errors:
            message = error.get('message', '').lower()
            categorised = False

            for category, pattern in patterns.items():
                if re.search(pattern, message):
                    categories[category] += 1
                    categorised = True
                    break

            if not categorised:
                categories['other'] += 1

        return categories

    def _detect_anomalies(self, errors: List, warnings: List) -> List[Dict]:
        """Detect anomalous log patterns."""
        anomalies = []

        # Spike detection: sudden increase in errors
        time_windows = self._group_by_time_window(errors, window_minutes=1)
        baseline_rate = sum(len(w) for w in time_windows[:5]) / 5  # First 5 minutes

        for i, window in enumerate(time_windows[5:], start=5):
            if len(window) > baseline_rate * 3:  # 3x increase
                anomalies.append({
                    'type': 'error_spike',
                    'time': i,
                    'count': len(window),
                    'baseline': baseline_rate
                })

        return anomalies

    def _query_range(self, query: str, start: datetime, end: datetime) -> List:
        """Execute Loki range query."""
        params = {
            'query': query,
            'start': int(start.timestamp() * 1000000000),
            'end': int(end.timestamp() * 1000000000),
            'limit': 5000
        }

        response = requests.get(f"{self.loki_url}/loki/api/v1/query_range", params=params)
        results = response.json()['data']['result']

        logs = []
        for stream in results:
            for value in stream['values']:
                timestamp, message = value
                logs.append({
                    'timestamp': datetime.fromtimestamp(int(timestamp) / 1000000000),
                    'message': message,
                    'labels': stream['stream']
                })

        return logs

    def _group_by_time_window(self, logs: List, window_minutes: int) -> List[List]:
        """Group logs into time windows."""
        if not logs:
            return []

        windows = []
        current_window = []
        window_start = logs[0]['timestamp']
        window_duration = timedelta(minutes=window_minutes)

        for log in logs:
            if log['timestamp'] - window_start > window_duration:
                windows.append(current_window)
                current_window = [log]
                window_start = log['timestamp']
            else:
                current_window.append(log)

        if current_window:
            windows.append(current_window)

        return windows

    def _create_timeline(self, errors: List, warnings: List) -> List[Dict]:
        """Create timeline of significant events."""
        events = []

        for error in errors:
            events.append({
                'timestamp': error['timestamp'],
                'type': 'error',
                'message': error['message']
            })

        for warning in warnings:
            events.append({
                'timestamp': warning['timestamp'],
                'type': 'warning',
                'message': warning['message']
            })

        # Sort by timestamp
        events.sort(key=lambda x: x['timestamp'])

        return events

Experiment Report Generation

# Generate comprehensive chaos experiment report
from dataclasses import dataclass
from typing import List, Dict
import json

@dataclass
class ExperimentReport:
    experiment_name: str
    hypothesis: str
    start_time: datetime
    end_time: datetime
    blast_radius: Dict

    # Metrics
    baseline_metrics: Dict
    experiment_metrics: Dict
    recovery_metrics: Dict

    # Observability data
    trace_analysis: Dict
    log_analysis: Dict

    # Results
    hypothesis_validated: bool
    findings: List[str]
    action_items: List[str]

    def generate_report(self) -> str:
        """Generate markdown report."""

        duration = (self.end_time - self.start_time).total_seconds() / 60

        report = f"""# Chaos Experiment Report: {self.experiment_name}

## Experiment Details

**Date:** {self.start_time.strftime('%Y-%m-%d %H:%M UTC')}
**Duration:** {duration:.1f} minutes
**Blast Radius:** {json.dumps(self.blast_radius, indent=2)}

## Hypothesis

{self.hypothesis}

**Result:** {'✅ VALIDATED' if self.hypothesis_validated else '❌ INVALIDATED'}

## Metrics Comparison

### Request Rate
- **Baseline:** {self.baseline_metrics.get('request_rate', 'N/A')} req/s
- **During Experiment:** {self.experiment_metrics.get('request_rate', 'N/A')} req/s
- **After Recovery:** {self.recovery_metrics.get('request_rate', 'N/A')} req/s

### Error Rate
- **Baseline:** {self.baseline_metrics.get('error_rate', 0):.2%}
- **During Experiment:** {self.experiment_metrics.get('error_rate', 0):.2%}
- **After Recovery:** {self.recovery_metrics.get('error_rate', 0):.2%}

### Latency (P99)
- **Baseline:** {self.baseline_metrics.get('p99_latency', 'N/A')}ms
- **During Experiment:** {self.experiment_metrics.get('p99_latency', 'N/A')}ms
- **After Recovery:** {self.recovery_metrics.get('p99_latency', 'N/A')}ms

## Observability Analysis

### Distributed Tracing
- **Total Traces:** {self.trace_analysis.get('total_traces', 0)}
- **Error Count:** {self.trace_analysis.get('error_count', 0)}
- **Retry Events:** {self.trace_analysis.get('retry_count', 0)}
- **Circuit Breaker Activations:** {self.trace_analysis.get('circuit_breaker_open', 0)}

### Log Analysis
- **Total Errors:** {self.log_analysis.get('total_errors', 0)}
- **Total Warnings:** {self.log_analysis.get('total_warnings', 0)}
- **Error Categories:** {json.dumps(self.log_analysis.get('error_categories', {}), indent=2)}

## Key Findings

"""
        for i, finding in enumerate(self.findings, 1):
            report += f"{i}. {finding}\n"

        report += "\n## Action Items\n\n"
        for i, action in enumerate(self.action_items, 1):
            report += f"- [ ] {action}\n"

        return report

    def save_to_file(self, filename: str):
        """Save report to markdown file."""
        with open(filename, 'w') as f:
            f.write(self.generate_report())

# Usage example
report = ExperimentReport(
    experiment_name="Payment Service Pod Failure",
    hypothesis="When 30% of payment-service pods fail, Istio retry logic maintains > 99.5% success rate",
    start_time=datetime(2024, 1, 15, 10, 0),
    end_time=datetime(2024, 1, 15, 10, 15),
    blast_radius={
        "environment": "production",
        "namespace": "payments",
        "affected_pods": "30%",
        "duration": "5 minutes"
    },
    baseline_metrics={
        "request_rate": 500,
        "error_rate": 0.001,
        "p99_latency": 180
    },
    experiment_metrics={
        "request_rate": 485,
        "error_rate": 0.003,
        "p99_latency": 220
    },
    recovery_metrics={
        "request_rate": 498,
        "error_rate": 0.001,
        "p99_latency": 185
    },
    trace_analysis={
        "total_traces": 15000,
        "error_count": 45,
        "retry_count": 120,
        "circuit_breaker_open": 0
    },
    log_analysis={
        "total_errors": 42,
        "total_warnings": 156,
        "error_categories": {
            "timeout": 5,
            "connection_refused": 37,
            "other": 0
        }
    },
    hypothesis_validated=True,
    findings=[
        "Istio retry logic successfully handled pod failures with minimal user impact",
        "Error rate remained below 0.5% throughout experiment",
        "P99 latency increased by 22% but stayed within acceptable bounds",
        "System fully recovered within 2 minutes of fault injection ending"
    ],
    action_items=[
        "Tune Istio retry timeout from 3s to 2s to reduce latency impact",
        "Add connection pool monitoring to detect exhaustion earlier",
        "Automate this experiment to run weekly in production"
    ]
)

report.save_to_file("chaos-experiment-2024-01-15.md")

Quick Reference

Chaos Experiment Workflow

Phase Duration Actions Validation
Planning 1-2 days Define hypothesis, select blast radius, identify metrics Peer review, stakeholder approval
Preparation 30 min Set up monitoring, configure safeguards, run preflight checks Health checks passing, alerts configured
Baseline 15 min Collect steady-state metrics before experiment Metrics within normal range
Execution 5-30 min Inject fault, monitor continuously Safeguards active, abort ready
Recovery 15-30 min Stop fault, verify system returns to steady state All metrics back to baseline
Analysis 1-2 hours Compare baseline vs experiment, validate hypothesis Report generated
Follow-up 1 week Implement improvements, document learnings Action items tracked

Common Experiment Types

Experiment Blast Radius Duration Key Metrics
Pod Failure 10-30% pods 2-5 min Error rate, replica recovery time
Network Latency 25-50% traffic 5-15 min P99 latency, timeout rate
CPU Stress 20-40% pods 10-20 min Response time, autoscaling trigger
Memory Pressure 10-20% pods 5-10 min OOM kills, pod restart rate
Zone Failure 1 AZ (of 3+) 10-30 min Cross-AZ traffic, availability
Dependency Failure 100% of dependency 3-10 min Circuit breaker activation, cache hit rate

Tool Selection Guide

Requirement Recommended Tool Reason
Kubernetes-only Chaos Mesh Native K8s integration, comprehensive
GitOps workflow Litmus Excellent CI/CD integration
Compliance/audit Gremlin Enterprise features, detailed reporting
AWS infrastructure AWS FIS Native AWS service integration
Multi-cloud Gremlin or Litmus Platform-agnostic
Cost-sensitive Chaos Mesh or Litmus Open source, no licensing

Safety Checklist

  • [ ] Hypothesis clearly defined
  • [ ] Blast radius limited and documented
  • [ ] Preflight health checks configured
  • [ ] Abort conditions specified
  • [ ] Manual abort procedure documented
  • [ ] Monitoring and alerting active
  • [ ] On-call engineer aware and available
  • [ ] Stakeholders notified
  • [ ] Rollback plan prepared
  • [ ] Postflight validation defined

Common Issues and Solutions

Issue: Experiment Causes Cascading Failures

Symptoms:

  • Blast radius expands beyond intended scope
  • Multiple services affected
  • Error rate exceeds abort thresholds
  • Recovery takes longer than expected

Solutions:

  • Implement circuit breakers on all service dependencies
    # Istio circuit breaker
    trafficPolicy:
      outlierDetection:
        consecutiveErrors: 5
        interval: 30s
        baseEjectionTime: 30s
    
  • Start with smaller blast radius (5-10% instead of 30%)
  • Add dependency health checks to abort conditions
  • Use bulkheads to isolate failure domains
  • Increase resource limits temporarily during experiments

Issue: Unable to Validate Hypothesis

Symptoms:

  • Results inconclusive
  • Metrics don't show expected behaviour
  • Missing observability data
  • Hypothesis too vague

Solutions:

  • Define specific, measurable hypothesis
    • Bad: "Service should be resilient"
    • Good: "Error rate < 1% when 30% pods fail for 5 minutes"
  • Add instrumentation before experimenting
    • Ensure traces include trace_id in logs
    • Add custom metrics for business logic
    • Configure log sampling appropriately
  • Extend baseline collection period (30 min instead of 15 min)
  • Run multiple iterations to account for variability

Issue: Production Experiments Too Risky

Symptoms:

  • Stakeholders uncomfortable with production chaos
  • Previous experiments caused incidents
  • Lack of confidence in abort mechanisms
  • Insufficient observability

Solutions:

  • Progress through environments gradually
    1. Local development
    2. Integration environment
    3. Staging with production-like load
    4. Production with minimal blast radius
    5. Expand production scope gradually
  • Implement feature flags for experiment control
    if feature_flag.is_enabled('chaos_experiments', user_id):
        # Apply chaos only to flagged traffic
    
  • Use canary releases to limit experiment scope
  • Run experiments during low-traffic periods initially
  • Require explicit approval for production experiments
  • Test abort mechanisms thoroughly in staging first

Issue: Experiments Don't Find Weaknesses

Symptoms:

  • All experiments pass
  • No improvements identified
  • Team confidence may be false
  • Not testing realistic scenarios

Solutions:

  • Increase experiment severity
    • Higher blast radius (30% → 50%)
    • Longer duration (5 min → 15 min)
    • Multiple simultaneous failures
  • Test more realistic scenarios
    # Combined network + resource stress
    apiVersion: chaos-mesh.org/v1alpha1
    kind: Workflow
    metadata:
      name: realistic-failure
    spec:
      entry: combined-chaos
      templates:
        - name: combined-chaos
          templateType: Parallel
          children:
            - network-latency
            - cpu-stress
            - pod-failure
    
  • Challenge assumptions in hypothesis
    • "Service handles 30% pod failure" → "Service handles 70% pod failure"
  • Test rare but impactful scenarios
    • Database connection pool exhaustion
    • DNS resolution failures
    • Dependency cascade failures
  • Review incident history for real-world failure modes

Issue: Alert Noise During Experiments

Symptoms:

  • Too many alerts fire
  • On-call paged unnecessarily
  • Hard to distinguish experiment from real incidents
  • Alert fatigue

Solutions:

  • Tag experiment resources
    metadata:
      labels:
        chaos.enabled: "true"
        chaos.experiment: "pod-failure-2024-01-15"
    
  • Suppress alerts during experiments
    # Alertmanager inhibition
    inhibit_rules:
      - source_match:
          chaos_experiment: "active"
        target_match:
          severity: "warning"
        equal: ['namespace', 'service']
    
  • Create separate chaos alert channel
    route:
      routes:
        - match:
            chaos: "true"
          receiver: chaos-team
          continue: true  # Still send to normal channels
    
  • Communicate experiment schedule to team
  • Use experiment annotations in Grafana to mark timeframes

Issue: Game Days Are Chaotic and Unproductive

Symptoms:

  • Game day sessions lack structure
  • No clear objectives or success criteria
  • Team overwhelmed by multiple failures
  • Learnings not documented

Solutions:

  • Create structured game day runbook
    # Game Day Runbook
    
    ## Objectives
    1. Test incident response procedures
    2. Validate RTO/RPO for database failure
    3. Train new team members on escalation
    
    ## Scenario Timeline
    09:00 - Intro and scenario briefing
    09:15 - Inject database primary failure
    09:20 - Teams respond, IC coordinates
    09:45 - Validate recovery completed
    10:00 - Debrief and retrospective
    
    ## Success Criteria
    - Database failover completes in < 5 minutes
    - No data loss
    - All team members know their roles
    - Runbook accurate and helpful
    
  • Assign roles clearly (Incident Commander, responders, observers)
  • Start with one scenario before adding complexity
  • Have facilitator who isn't responding to incident
  • Schedule retrospective immediately after (while fresh)
  • Document action items and assign owners

Issue: Chaos Experiments Not Integrated into CI/CD

Symptoms:

  • Experiments run ad-hoc
  • No automated regression testing for resilience
  • Regressions introduced by new deployments
  • Manual effort to run experiments

Solutions:

  • Add chaos tests to CI pipeline
    # GitLab CI example
    chaos-test:
      stage: test
      script:
        - kubectl apply -f chaos/network-latency.yaml
        - ./scripts/wait-for-experiment.sh
        - ./scripts/validate-metrics.sh
        - kubectl delete -f chaos/network-latency.yaml
      only:
        - merge_requests
      environment:
        name: staging
    
  • Use Litmus with GitOps
    # Argo Workflow integration
    apiVersion: argoproj.io/v1alpha1
    kind: Workflow
    metadata:
      name: chaos-pipeline
    spec:
      entrypoint: run-chaos-tests
      templates:
        - name: run-chaos-tests
          steps:
            - - name: deploy-app
                template: deploy
            - - name: pod-failure-test
                template: litmus-experiment
                arguments:
                  parameters:
                    - name: experiment
                      value: "pod-delete"
            - - name: validate-results
                template: check-metrics
    
  • Schedule recurring experiments
    # CronJob for weekly chaos
    apiVersion: batch/v1
    kind: CronJob
    metadata:
      name: weekly-pod-chaos
    spec:
      schedule: "0 2 * * MON"  # Every Monday 2am
      jobTemplate:
        spec:
          template:
            spec:
              containers:
                - name: chaos-runner
                  image: chaostoolkit/chaostoolkit
                  command:
                    - chaos
                    - run
                    - /experiments/pod-failure.yaml
    
  • Fail builds if resilience tests fail

Related Topics

The following topics would complement this Chaos Engineering Practices cheatsheet:

  1. Observability Patterns - Comprehensive guide to metrics, logs, and traces for validating chaos experiments and measuring system health

  2. SLOs, SLIs, and Error Budgets - Detailed coverage of defining service level objectives and using error budgets to guide chaos engineering priorities

  3. Kubernetes Advanced - Deep dive into Kubernetes resilience features like pod disruption budgets, health checks, and graceful shutdown

  4. Istio/Service Mesh - Service mesh patterns for traffic management, circuit breaking, retries, and fault injection at the network level

  5. Incident Management and Postmortems - Best practices for responding to real incidents and conducting blameless postmortems to drive improvements

  6. Site Reliability Engineering (SRE) Practices - Broader SRE principles including toil reduction, capacity planning, and production readiness reviews that complement chaos engineering