Argo Rollouts
Progressive delivery controller for Kubernetes providing advanced deployment strategies like canary and blue/green.
Argo Rollouts
Progressive delivery controller for Kubernetes providing advanced deployment strategies like canary and blue/green.
Overview
Argo Rollouts is a Kubernetes controller that provides advanced deployment capabilities beyond standard Kubernetes Deployments. It enables progressive delivery through canary deployments, blue/green deployments, and automated analysis to minimise risk during application updates. Rollouts integrates with ingress controllers and service meshes to gradually shift traffic, and supports automated rollback based on metrics from Prometheus, Datadog, New Relic, and other providers.
flowchart TD
subgraph "Progressive Delivery Workflow"
A[New Version] --> B[Rollout Controller]
B --> C{Strategy?}
C -->|Canary| D[Gradual Traffic Shift]
C -->|Blue/Green| E[Full Switch]
D --> F[Analysis]
E --> F
F --> G{Metrics OK?}
G -->|Yes| H[Continue/Complete]
G -->|No| I[Automatic Rollback]
H --> J[100% New Version]
I --> K[100% Old Version]
end
Installation and Setup
Argo Rollouts consists of a controller that watches Rollout resources and manages progressive delivery.
Key Concepts
- Controller: Manages Rollout resources and orchestrates deployments
- Kubectl Plugin: CLI tool for managing and monitoring rollouts
- Dashboard: Optional UI for visualising rollout status
- CRDs: Custom Resource Definitions for Rollout, AnalysisTemplate, AnalysisRun, Experiment
graph TB
subgraph "Argo Rollouts Architecture"
A[Rollout Controller] --> B[Rollout CRD]
A --> C[AnalysisRun]
A --> D[ReplicaSets]
B --> E[Strategy Config]
B --> F[Traffic Management]
C --> G[Metrics Providers]
F --> H[Ingress/Service Mesh]
D --> I[Pods]
end
Installation Commands
# Install Argo Rollouts controller
kubectl create namespace argo-rollouts
kubectl apply -n argo-rollouts -f https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml
# Install via Helm
helm repo add argo https://argoproj.github.io/argo-helm
helm install argo-rollouts argo/argo-rollouts -n argo-rollouts --create-namespace
# Install kubectl plugin (Linux)
curl -LO https://github.com/argoproj/argo-rollouts/releases/latest/download/kubectl-argo-rollouts-linux-amd64
chmod +x kubectl-argo-rollouts-linux-amd64
sudo mv kubectl-argo-rollouts-linux-amd64 /usr/local/bin/kubectl-argo-rollouts
# Install kubectl plugin (macOS)
brew install argoproj/tap/kubectl-argo-rollouts
# Verify installation
kubectl argo rollouts version
# Install dashboard (optional)
kubectl apply -n argo-rollouts -f https://github.com/argoproj/argo-rollouts/releases/latest/download/dashboard-install.yaml
# Access dashboard
kubectl port-forward -n argo-rollouts svc/argo-rollouts-dashboard 3100:3100
# Visit http://localhost:3100
Controller Configuration
apiVersion: apps/v1
kind: Deployment
metadata:
name: argo-rollouts
namespace: argo-rollouts
spec:
replicas: 2
selector:
matchLabels:
app.kubernetes.io/name: argo-rollouts
template:
metadata:
labels:
app.kubernetes.io/name: argo-rollouts
spec:
serviceAccountName: argo-rollouts
containers:
- name: argo-rollouts
image: quay.io/argoproj/argo-rollouts:stable
args:
- --loglevel
- info
- --leader-elect
- --namespaced # Watch specific namespaces
ports:
- containerPort: 8090
name: metrics
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
capabilities:
drop:
- ALL
Rollout Objects and Strategies
Rollouts are Kubernetes custom resources that extend Deployment functionality with progressive delivery strategies.
Key Concepts
- Rollout: Similar to Deployment but with advanced deployment strategies
- Strategy: Defines how updates are rolled out (Canary, Blue/Green)
- Revision: Each update creates a new revision tracked by the controller
- Stable/Canary ReplicaSets: Managed automatically during progressive delivery
- Pod Template Hash: Label used to track versions
flowchart LR
subgraph "Rollout Structure"
A[Rollout] --> B[Strategy]
A --> C[Pod Template]
A --> D[Replicas]
B --> E[Canary Steps]
B --> F[Blue/Green Config]
B --> G[Analysis]
end
Basic Rollout Manifest
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp
namespace: production
spec:
replicas: 5
# Revision history to keep
revisionHistoryLimit: 3
# Selector must match pod template labels
selector:
matchLabels:
app: myapp
# Pod template (same as Deployment)
template:
metadata:
labels:
app: myapp
version: stable
spec:
containers:
- name: myapp
image: myapp:v1.0.0
ports:
- containerPort: 8080
name: http
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
# Strategy defined separately (see following sections)
strategy:
canary: {}
Common Commands
# List rollouts
kubectl argo rollouts list -n production
# Get rollout status
kubectl argo rollouts get rollout myapp -n production
# Watch rollout progress
kubectl argo rollouts get rollout myapp -n production --watch
# Describe rollout
kubectl describe rollout myapp -n production
# View rollout history
kubectl argo rollouts history rollout myapp -n production
# Trigger a rollout by updating image
kubectl argo rollouts set image myapp myapp=myapp:v2.0.0 -n production
# Manually promote a rollout
kubectl argo rollouts promote myapp -n production
# Skip all remaining steps and deploy immediately
kubectl argo rollouts promote myapp -n production --full
# Pause a rollout
kubectl argo rollouts pause myapp -n production
# Resume a paused rollout
kubectl argo rollouts resume myapp -n production
# Abort a rollout
kubectl argo rollouts abort myapp -n production
# Retry a failed rollout
kubectl argo rollouts retry rollout myapp -n production
# Undo to previous revision
kubectl argo rollouts undo myapp -n production
# Undo to specific revision
kubectl argo rollouts undo myapp --to-revision=3 -n production
# Restart rollout (create new revision)
kubectl argo rollouts restart myapp -n production
# Get rollout status (non-interactive)
kubectl argo rollouts status myapp -n production
Converting Deployment to Rollout
# Before (Deployment)
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp
spec:
replicas: 5
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v1.0.0
---
# After (Rollout)
apiVersion: argoproj.io/v1alpha1
kind: Rollout # Changed from Deployment
metadata:
name: myapp
spec:
replicas: 5
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v1.0.0
strategy: # Added strategy
canary:
steps:
- setWeight: 20
- pause: {duration: 1m}
- setWeight: 50
- pause: {duration: 2m}
Canary Deployment Strategy
Canary deployments gradually shift traffic to the new version in controlled steps, allowing validation at each stage.
Key Concepts
- Traffic Weight: Percentage of traffic sent to canary version
- Steps: Sequence of weight changes and pauses
- Manual/Auto Promotion: Whether steps advance automatically or require approval
- MaxSurge/MaxUnavailable: Control pod creation during rollout
- Canary Service: Service pointing to canary pods for testing
sequenceDiagram
participant User
participant Controller as Rollout Controller
participant Stable as Stable ReplicaSet
participant Canary as Canary ReplicaSet
participant Analysis as Analysis
User->>Controller: Update image
Controller->>Canary: Create v2 (10% weight)
Controller->>Analysis: Start metrics check
Analysis-->>Controller: Metrics pass
Controller->>Canary: Scale to 25% weight
Controller->>Analysis: Continue monitoring
Analysis-->>Controller: Metrics pass
Controller->>Canary: Scale to 50% weight
Note over Controller: Manual gate - wait for promotion
User->>Controller: Promote
Controller->>Canary: Scale to 100% weight
Controller->>Stable: Scale down v1
Controller-->>User: Rollout complete
Canary Strategy Configuration
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-canary
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v1.0.0
ports:
- containerPort: 8080
strategy:
canary:
# Maximum additional replicas during rollout
maxSurge: "25%"
# Maximum unavailable replicas during rollout
maxUnavailable: 0
# Canary steps
steps:
# Step 1: Send 10% of traffic to canary
- setWeight: 10
- pause: {duration: 5m}
# Step 2: Increase to 25%
- setWeight: 25
- pause: {duration: 5m}
# Step 3: Increase to 50%
- setWeight: 50
- pause: {duration: 10m}
# Step 4: Increase to 75%
- setWeight: 75
- pause: {duration: 10m}
# Step 5: Manual promotion pause (wait indefinitely)
- pause: {}
# Final: 100% to new version (implicit)
# Optional: Analysis at specific steps
analysis:
templates:
- templateName: success-rate
startingStep: 2 # Start analysis at step index 2
# Traffic routing (see integration section)
trafficRouting:
nginx:
stableIngress: myapp-stable
annotationPrefix: nginx.ingress.kubernetes.io
Canary with Background Analysis
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-auto-analysis
spec:
replicas: 5
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v1.0.0
strategy:
canary:
steps:
- setWeight: 20
- pause: {duration: 2m}
- setWeight: 40
- pause: {duration: 2m}
- setWeight: 60
- pause: {duration: 2m}
- setWeight: 80
- pause: {duration: 2m}
# Background analysis runs throughout the rollout
analysis:
templates:
- templateName: error-rate
- templateName: response-time
# Analysis runs for entire canary deployment
args:
- name: service-name
value: myapp
- name: canary-hash
valueFrom:
podTemplateHashValue: Latest
Canary with Manual Gates
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: production-app
spec:
replicas: 10
selector:
matchLabels:
app: production-app
template:
metadata:
labels:
app: production-app
spec:
containers:
- name: app
image: production-app:v2.0.0
strategy:
canary:
maxSurge: "20%"
maxUnavailable: 0
steps:
# Initial canary with 5%
- setWeight: 5
- pause: {duration: 5m}
# Expand to 10% with manual approval
- setWeight: 10
- pause: {} # Wait indefinitely for promotion
# Progressive rollout with analysis
- setWeight: 25
- pause: {duration: 10m}
- setWeight: 50
- pause: {} # Another manual gate
- setWeight: 75
- pause: {duration: 10m}
# Final manual approval before 100%
- pause: {}
analysis:
templates:
- templateName: critical-metrics
startingStep: 2
Blue/Green Deployment Strategy
Blue/green deployments maintain two complete environments, switching traffic atomically between them.
Key Concepts
- Active Service: Service pointing to stable (blue) version
- Preview Service: Service pointing to new (green) version for testing
- Automatic Promotion: Can automatically promote after duration
- Instant Switchover: Traffic switches immediately, no gradual shift
- Preview Testing: Test new version before promoting to active
flowchart TB
subgraph "Blue/Green Deployment Flow"
A[Current: Blue v1] --> B[Deploy Green v2]
B --> C[Preview Service<br/>points to Green]
C --> D{Run Pre-Promotion<br/>Analysis}
D -->|Pass| E[Switch Active Service<br/>to Green]
D -->|Fail| F[Abort - Stay on Blue]
E --> G{Run Post-Promotion<br/>Analysis}
G -->|Pass| H[Scale Down Blue]
G -->|Fail| I[Rollback to Blue]
H --> J[Complete]
I --> K[Blue Still Active]
end
Blue/Green Strategy Configuration
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-bluegreen
spec:
replicas: 5
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v1.0.0
ports:
- containerPort: 8080
strategy:
blueGreen:
# Active service (production traffic)
activeService: myapp-active
# Preview service (for testing new version)
previewService: myapp-preview
# Automatically promote after duration
autoPromotionEnabled: false
autoPromotionSeconds: 300 # 5 minutes
# Scale down old version after promotion
scaleDownDelaySeconds: 30
scaleDownDelayRevisionLimit: 2
# Anti-affinity between blue and green
antiAffinity:
requiredDuringSchedulingIgnoredDuringExecution: {}
# Optional: Run analysis before promotion
prePromotionAnalysis:
templates:
- templateName: smoke-tests
args:
- name: service-url
value: http://myapp-preview
# Optional: Run analysis after promotion
postPromotionAnalysis:
templates:
- templateName: integration-tests
args:
- name: service-url
value: http://myapp-active
Required Services for Blue/Green
# Active service - receives production traffic
apiVersion: v1
kind: Service
metadata:
name: myapp-active
spec:
selector:
app: myapp
# Rollout controller manages this label automatically
ports:
- port: 80
targetPort: 8080
protocol: TCP
type: ClusterIP
---
# Preview service - for testing new version
apiVersion: v1
kind: Service
metadata:
name: myapp-preview
spec:
selector:
app: myapp
# Rollout controller manages this label automatically
ports:
- port: 80
targetPort: 8080
protocol: TCP
type: ClusterIP
Blue/Green with Comprehensive Analysis
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: critical-app
spec:
replicas: 10
selector:
matchLabels:
app: critical-app
template:
metadata:
labels:
app: critical-app
spec:
containers:
- name: app
image: critical-app:v2.0.0
ports:
- containerPort: 8080
strategy:
blueGreen:
activeService: critical-app-active
previewService: critical-app-preview
# Manual promotion only
autoPromotionEnabled: false
# Keep old version for 1 hour after promotion
scaleDownDelaySeconds: 3600
scaleDownDelayRevisionLimit: 1
# Pre-promotion smoke tests
prePromotionAnalysis:
templates:
- templateName: smoke-tests
- templateName: load-tests
args:
- name: preview-url
value: http://critical-app-preview
- name: duration
value: "300"
# Post-promotion verification
postPromotionAnalysis:
templates:
- templateName: error-rate-check
- templateName: performance-check
args:
- name: active-url
value: http://critical-app-active
- name: threshold
value: "0.01" # 1% error rate
Canary vs Blue/Green Comparison
Understanding when to use each strategy based on your requirements.
Strategy Comparison
| Aspect | Canary | Blue/Green |
|---|---|---|
| Traffic Shift | Gradual (percentage-based) | Instant (atomic switch) |
| Resource Usage | Efficient (gradual scaling) | Higher (2x pods during switch) |
| Rollback Speed | Immediate (shift back) | Immediate (switch back) |
| Testing Ability | Production traffic sampling | Full preview environment |
| Complexity | Higher (traffic management) | Lower (simple switch) |
| Risk | Lower (gradual exposure) | Higher (instant full exposure) |
| Best For | Production deployments | Staging, critical systems |
| User Impact | Gradual | Immediate |
| Infrastructure | Needs ingress/mesh | Standard services |
| Validation | Real user traffic | Isolated testing |
| Granularity | Fine-grained control | All-or-nothing |
graph TB
subgraph "Canary Deployment"
A1[100% v1] --> B1[90% v1, 10% v2]
B1 --> C1[75% v1, 25% v2]
C1 --> D1[50% v1, 50% v2]
D1 --> E1[25% v1, 75% v2]
E1 --> F1[100% v2]
end
subgraph "Blue/Green Deployment"
A2[100% Blue v1] --> B2[Test Green v2<br/>via Preview]
B2 --> C2[Switch: 100% Green v2]
C2 --> D2[Scale down Blue]
end
Decision Matrix
flowchart TD
A[Choose Deployment Strategy] --> B{Need preview<br/>environment?}
B -->|Yes| C[Blue/Green]
B -->|No| D{Risk tolerance?}
D -->|Low risk tolerance| E{Have ingress/<br/>service mesh?}
E -->|Yes| F[Canary]
E -->|No| C
D -->|High risk tolerance| G{Need gradual<br/>rollout?}
G -->|Yes| F
G -->|No| C
F --> H[Canary Strategy]
C --> I[Blue/Green Strategy]
Use Case Recommendations
Use Canary When:
- Deploying to production with real user traffic
- Want to limit blast radius of potential issues
- Have metrics to validate new version
- Can integrate with ingress controller or service mesh
- Need fine-grained control over traffic shift
- Cost-conscious (gradual scaling more efficient)
- Want to validate with real traffic patterns
- Need to detect issues early with minimal impact
Use Blue/Green When:
- Need isolated preview environment for testing
- Require instant cutover (e.g., coordinated releases)
- Don't have advanced traffic routing available
- Deploy to staging/pre-production environments
- Need simple, deterministic deployment process
- Can afford temporary resource doubling
- Require quick, complete rollback capability
- Want to run comprehensive tests before cutover
- Database schema changes require atomic switch
Hybrid Approach
# Staging: Blue/Green for quick testing
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-staging
namespace: staging
spec:
replicas: 3
selector:
matchLabels:
app: myapp
env: staging
template:
metadata:
labels:
app: myapp
env: staging
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
blueGreen:
activeService: myapp-staging
previewService: myapp-staging-preview
autoPromotionEnabled: true
autoPromotionSeconds: 60
prePromotionAnalysis:
templates:
- templateName: basic-checks
---
# Production: Canary for careful rollout
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-production
namespace: production
spec:
replicas: 20
selector:
matchLabels:
app: myapp
env: production
template:
metadata:
labels:
app: myapp
env: production
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
maxSurge: "25%"
maxUnavailable: 0
steps:
- setWeight: 5
- pause: {duration: 10m}
- setWeight: 20
- pause: {duration: 10m}
- setWeight: 50
- pause: {} # Manual gate
- setWeight: 80
- pause: {duration: 15m}
analysis:
templates:
- templateName: error-rate
- templateName: latency
startingStep: 1
Analysis Templates and Metrics Providers
Analysis templates define success criteria for deployments using metrics from various providers.
Key Concepts
- AnalysisTemplate: Reusable metric query and success criteria definition
- AnalysisRun: Instance of analysis execution during rollout
- Metrics Provider: Backend system providing metrics (Prometheus, Datadog, etc.)
- Success Criteria: Thresholds determining pass/fail
- Arguments: Parameters passed from Rollout to template
flowchart TB
subgraph "Analysis Workflow"
A[Rollout Triggers] --> B[Create AnalysisRun]
B --> C[AnalysisTemplate]
C --> D[Define Metrics]
D --> E[Query Provider]
E --> F{Evaluate<br/>Criteria}
F -->|Success| G[Continue Rollout]
F -->|Failure| H[Abort & Rollback]
F -->|Error| I[Inconclusive]
I --> J[Retry or Fail]
end
AnalysisTemplate Structure
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: success-rate
namespace: production
spec:
# Arguments passed from Rollout
args:
- name: service-name
- name: namespace
value: production # Default value
# Metrics to evaluate
metrics:
- name: success-rate
# How long to run analysis
interval: 1m
# Number of measurements
count: 5
# Maximum consecutive failures before abort
failureLimit: 2
# Initial delay before starting
initialDelay: 30s
# Success criteria
successCondition: result >= 0.95
failureCondition: result < 0.90
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
sum(rate(http_requests_total{
service="{{args.service-name}}",
namespace="{{args.namespace}}",
status=~"2.."
}[5m]))
/
sum(rate(http_requests_total{
service="{{args.service-name}}",
namespace="{{args.namespace}}"
}[5m]))
Prometheus Provider
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: prometheus-metrics
spec:
args:
- name: service
- name: threshold
value: "0.99"
metrics:
# Error rate check
- name: error-rate
interval: 30s
count: 10
successCondition: result < 0.01 # Less than 1% error rate
failureCondition: result >= 0.05 # More than 5% error rate
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
sum(rate(http_requests_total{
service="{{args.service}}",
status=~"5.."
}[2m]))
/
sum(rate(http_requests_total{
service="{{args.service}}"
}[2m]))
# Response time check (p95)
- name: response-time-p95
interval: 30s
count: 10
successCondition: result < 500 # Less than 500ms
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket{
service="{{args.service}}"
}[2m])) by (le)
) * 1000
# CPU usage check
- name: cpu-usage
interval: 1m
count: 5
successCondition: result < 80 # Less than 80%
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
sum(rate(container_cpu_usage_seconds_total{
pod=~"{{args.service}}.*"
}[2m])) by (pod)
/
sum(container_spec_cpu_quota{
pod=~"{{args.service}}.*"
} / container_spec_cpu_period{
pod=~"{{args.service}}.*"
}) by (pod)
* 100
# Memory usage check
- name: memory-usage
interval: 1m
count: 5
successCondition: result < 85 # Less than 85%
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
sum(container_memory_working_set_bytes{
pod=~"{{args.service}}.*"
}) by (pod)
/
sum(container_spec_memory_limit_bytes{
pod=~"{{args.service}}.*"
}) by (pod)
* 100
Datadog Provider
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: datadog-metrics
spec:
args:
- name: service-name
metrics:
- name: error-rate
interval: 1m
count: 5
successCondition: result < 5
provider:
datadog:
apiVersion: v1
interval: 5m
query: |
sum:trace.http.request.errors{
service:{{args.service-name}}
}.as_count()
/
sum:trace.http.request.hits{
service:{{args.service-name}}
}.as_count()
* 100
- name: avg-response-time
interval: 1m
count: 5
successCondition: result < 0.5 # 500ms
provider:
datadog:
apiVersion: v1
interval: 5m
query: |
avg:trace.http.request.duration{
service:{{args.service-name}}
}
- name: apm-errors
interval: 2m
count: 5
successCondition: result < 10
provider:
datadog:
apiVersion: v1
interval: 5m
query: |
sum:trace.errors{
service:{{args.service-name}}
}.as_count()
New Relic Provider
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: newrelic-metrics
spec:
args:
- name: service-name
- name: canary-hash
metrics:
- name: apdex-score
interval: 2m
count: 5
successCondition: result > 0.95
provider:
newRelic:
profile: production-account
query: |
FROM Transaction
SELECT apdex(duration, t: 0.5)
WHERE appName = '{{args.service-name}}'
AND tags.revision = '{{args.canary-hash}}'
- name: error-percentage
interval: 2m
count: 5
successCondition: result < 1
provider:
newRelic:
profile: production-account
query: |
FROM Transaction
SELECT percentage(count(*), WHERE error IS true)
WHERE appName = '{{args.service-name}}'
AND tags.revision = '{{args.canary-hash}}'
- name: throughput
interval: 2m
count: 5
successCondition: result > 100
provider:
newRelic:
profile: production-account
query: |
FROM Transaction
SELECT rate(count(*), 1 minute)
WHERE appName = '{{args.service-name}}'
AND tags.revision = '{{args.canary-hash}}'
Web/HTTP Provider
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: http-health-check
spec:
args:
- name: service-url
metrics:
- name: health-check
interval: 30s
count: 10
successCondition: result == "200"
failureLimit: 3
provider:
web:
url: "{{args.service-url}}/health"
headers:
- key: User-Agent
value: argo-rollouts
jsonPath: "{$.status}" # Parse JSON response
timeoutSeconds: 10
- name: api-response-check
interval: 1m
count: 5
successCondition: "result.healthy == true"
provider:
web:
url: "{{args.service-url}}/api/status"
jsonPath: "{$}"
- name: metrics-endpoint
interval: 1m
count: 5
successCondition: result.error_rate < 0.01
provider:
web:
url: "{{args.service-url}}/metrics"
jsonPath: "{$.error_rate}"
Job Provider
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: load-test
spec:
args:
- name: target-url
metrics:
- name: load-test-results
provider:
job:
spec:
backoffLimit: 0
template:
spec:
containers:
- name: load-test
image: grafana/k6:latest
command:
- k6
- run
- /scripts/load-test.js
env:
- name: TARGET_URL
value: "{{args.target-url}}"
volumeMounts:
- name: script
mountPath: /scripts
volumes:
- name: script
configMap:
name: k6-load-test-script
restartPolicy: Never
---
apiVersion: v1
kind: ConfigMap
metadata:
name: k6-load-test-script
data:
load-test.js: |
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
stages: [
{ duration: '1m', target: 50 },
{ duration: '2m', target: 50 },
{ duration: '1m', target: 0 },
],
thresholds: {
http_req_duration: ['p(95)<500'],
http_req_failed: ['rate<0.01'],
},
};
export default function () {
const res = http.get(__ENV.TARGET_URL);
check(res, {
'status is 200': (r) => r.status === 200,
});
sleep(1);
}
Combining Multiple Templates
# Rollout using multiple analysis templates
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: comprehensive-analysis
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
steps:
- setWeight: 20
- pause: {duration: 5m}
- setWeight: 50
- pause: {duration: 5m}
analysis:
templates:
# Combine multiple templates
- templateName: prometheus-metrics
- templateName: datadog-metrics
- templateName: http-health-check
args:
- name: service
value: myapp
- name: service-name
value: myapp
- name: service-url
value: http://myapp-canary
Multi-Provider Analysis
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: comprehensive-checks
spec:
args:
- name: service
- name: canary-hash
metrics:
# Prometheus: Error rate
- name: error-rate
interval: 2m
count: 5
successCondition: result < 0.01
failureLimit: 2
provider:
prometheus:
address: http://prometheus:9090
query: |
sum(rate(http_requests_total{
service="{{args.service}}",
rollouts_pod_template_hash="{{args.canary-hash}}",
status=~"5.."
}[5m]))
/
sum(rate(http_requests_total{
service="{{args.service}}",
rollouts_pod_template_hash="{{args.canary-hash}}"
}[5m]))
# New Relic: Apdex score
- name: apdex
interval: 2m
count: 5
successCondition: result > 0.95
provider:
newRelic:
profile: my-account
query: |
FROM Transaction
SELECT apdex(duration, t: 0.5)
WHERE appName = '{{args.service}}'
AND tags.revision = '{{args.canary-hash}}'
# Datadog: Log errors
- name: log-errors
interval: 2m
count: 5
successCondition: result < 10
provider:
datadog:
interval: 5m
query: |
sum:log.error.count{
service:{{args.service}},
revision:{{args.canary-hash}}
}.as_count()
Progressive Delivery Steps and Pauses
Controlling the pace and gates of rollout progression through steps and pauses.
Key Concepts
- Steps: Ordered sequence of actions during canary rollout
- setWeight: Traffic percentage to send to canary
- pause: Wait period (duration or indefinite)
- experiment: A/B testing with multiple variants
- analysis: Metric evaluation at specific steps
- setCanaryScale: Control number of canary replicas independently
stateDiagram-v2
[*] --> Step1: Start Rollout
Step1 --> Pause1: setWeight 10%
Pause1 --> Analysis1: Duration elapsed
Analysis1 --> Step2: Metrics pass
Analysis1 --> Rollback: Metrics fail
Step2 --> Pause2: setWeight 50%
Pause2 --> ManualGate: Duration elapsed
ManualGate --> Step3: Manual promote
ManualGate --> Rollback: Manual abort
Step3 --> Complete: setWeight 100%
Rollback --> [*]: Return to stable
Complete --> [*]
Step Types
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: step-examples
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
steps:
# Set traffic weight
- setWeight: 10
# Pause for specific duration
- pause:
duration: 5m
# Pause indefinitely (manual promotion required)
- pause: {}
# Set traffic weight with canary replicas
- setWeight: 30
- setCanaryScale:
weight: 30 # Scale canary to 30% of spec.replicas
# Or set explicit replica count
- setCanaryScale:
replicas: 3 # Exactly 3 canary replicas
# Inline analysis
- analysis:
templates:
- templateName: success-rate
args:
- name: service-name
value: myapp
# Continue progression
- setWeight: 50
- pause:
duration: 10m
# Experiment (A/B testing)
- experiment:
duration: 10m
templates:
- name: experiment-baseline
specRef: stable
weight: 50
- name: experiment-canary
specRef: canary
weight: 50
analyses:
- name: experiment-analysis
templateName: experiment-comparison
# Final steps
- setWeight: 100
Dynamic Pauses
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: dynamic-pause
spec:
replicas: 5
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
steps:
- setWeight: 20
# Automatic pause (analysis-driven)
- pause:
duration: 5m
# Progress to next step automatically if analysis passes
- analysis:
templates:
- templateName: quick-check
- setWeight: 50
# Manual gate - requires explicit promotion
- pause: {}
# Note: Manual promotion command
# kubectl argo rollouts promote myapp
Anti-Flapping Configuration
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: anti-flapping
spec:
replicas: 10
# Minimum time before health check
minReadySeconds: 30
# Maximum time to wait for rollout to progress
progressDeadlineSeconds: 600
# Abort on progress deadline
progressDeadlineAbort: true
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
maxSurge: 1
maxUnavailable: 0
steps:
- setWeight: 25
- pause:
duration: 2m
- setWeight: 50
- pause:
duration: 2m
- setWeight: 75
- pause:
duration: 2m
Background Analysis Throughout Rollout
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: background-analysis
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
steps:
- setWeight: 20
- pause: {duration: 5m}
- setWeight: 40
- pause: {duration: 5m}
- setWeight: 60
- pause: {duration: 5m}
- setWeight: 80
- pause: {duration: 5m}
# Background analysis runs continuously
analysis:
templates:
- templateName: continuous-monitoring
# Start analysis at step index 1
startingStep: 1
args:
- name: service-name
value: myapp
- name: canary-hash
valueFrom:
podTemplateHashValue: Latest
Multi-Stage Analysis
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: multi-stage-analysis
spec:
replicas: 20
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
maxSurge: "25%"
maxUnavailable: 0
steps:
# Stage 1: Initial rollout with basic health checks
- setWeight: 5
- pause: {duration: 2m}
- analysis:
templates:
- templateName: basic-health
# Stage 2: Expand with error rate monitoring
- setWeight: 20
- pause: {duration: 5m}
- analysis:
templates:
- templateName: error-rate-check
- templateName: latency-check
# Stage 3: Half traffic with comprehensive analysis
- setWeight: 50
- pause: {duration: 10m}
- analysis:
templates:
- templateName: error-rate-check
- templateName: latency-check
- templateName: business-metrics
# Stage 4: Manual approval gate
- pause: {}
# Stage 5: Final rollout
- setWeight: 80
- pause: {duration: 10m}
- analysis:
templates:
- templateName: final-verification
# Background analysis throughout
analysis:
templates:
- templateName: continuous-alerts
startingStep: 1
Automated Rollback Scenarios
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: auto-rollback
spec:
replicas: 10
# Abort if rollout doesn't progress
progressDeadlineSeconds: 600
progressDeadlineAbort: true
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
maxSurge: "25%"
maxUnavailable: 0
# Automatically abort on failed analysis
abortScaleDownDelaySeconds: 30
steps:
- setWeight: 10
- pause: {duration: 2m}
# Analysis will auto-rollback if metrics fail
- analysis:
templates:
- templateName: critical-metrics
args:
- name: service
value: myapp
- name: canary-hash
valueFrom:
podTemplateHashValue: Latest
- setWeight: 30
- pause: {duration: 5m}
# Multiple analysis templates - ANY failure triggers rollback
- analysis:
templates:
- templateName: error-rate
- templateName: latency
- templateName: saturation
args:
- name: service
value: myapp
- setWeight: 50
- pause: {}
---
# Analysis template with strict failure conditions
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: critical-metrics
spec:
args:
- name: service
- name: canary-hash
metrics:
- name: error-rate
interval: 1m
count: 5
# Allow only 1 consecutive failure
failureLimit: 1
# Inconclusive results treated as failures
inconclusiveLimit: 1
successCondition: result < 0.01
# Hard failure threshold
failureCondition: result >= 0.05
provider:
prometheus:
address: http://prometheus:9090
query: |
sum(rate(http_requests_total{
service="{{args.service}}",
rollouts_pod_template_hash="{{args.canary-hash}}",
status=~"5.."
}[5m]))
/
sum(rate(http_requests_total{
service="{{args.service}}",
rollouts_pod_template_hash="{{args.canary-hash}}"
}[5m]))
Integration with Ingress Controllers
Integrating Argo Rollouts with ingress controllers for traffic splitting.
Key Concepts
- Traffic Routing: Ingress controller routes traffic based on weights
- Stable/Canary Services: Separate services for each version
- Header-Based Routing: Route based on headers for testing
- Annotation Management: Rollout controller updates ingress annotations
flowchart TD
subgraph "Traffic Management"
A[Internet] --> B[Ingress Controller]
B --> C{Traffic Split}
C -->|80%| D[Stable Service]
C -->|20%| E[Canary Service]
D --> F[Stable Pods v1]
E --> G[Canary Pods v2]
H[Rollout Controller] --> B
H --> I[Update Weights]
end
NGINX Ingress
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-nginx
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
ports:
- containerPort: 8080
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
nginx:
# Name of main ingress
stableIngress: myapp-ingress
# Optional: Annotation prefix (default: nginx.ingress.kubernetes.io)
annotationPrefix: nginx.ingress.kubernetes.io
# Optional: Additional ingress annotations
additionalIngressAnnotations:
canary-by-header: X-Canary
canary-by-header-value: "true"
steps:
- setWeight: 10
- pause: {duration: 5m}
- setWeight: 25
- pause: {duration: 5m}
- setWeight: 50
- pause: {duration: 10m}
- setWeight: 75
- pause: {duration: 10m}
---
# Stable service
apiVersion: v1
kind: Service
metadata:
name: myapp-stable
spec:
selector:
app: myapp
ports:
- port: 80
targetPort: 8080
---
# Canary service
apiVersion: v1
kind: Service
metadata:
name: myapp-canary
spec:
selector:
app: myapp
ports:
- port: 80
targetPort: 8080
---
# Ingress
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: myapp-ingress
annotations:
nginx.ingress.kubernetes.io/rewrite-target: /
spec:
ingressClassName: nginx
rules:
- host: myapp.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: myapp-stable
port:
number: 80
ALB (AWS Application Load Balancer)
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-alb
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
ports:
- containerPort: 8080
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
alb:
# Name of the ingress
ingress: myapp-ingress
# Service port
servicePort: 80
# Optional: Root service (for multiple paths)
rootService: myapp-root
# Optional: Sticky sessions
stickinessConfig:
enabled: true
durationSeconds: 3600
steps:
- setWeight: 10
- pause: {duration: 5m}
- setWeight: 30
- pause: {duration: 10m}
- setWeight: 50
- pause: {}
---
# ALB Ingress
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: myapp-ingress
annotations:
alb.ingress.kubernetes.io/scheme: internet-facing
alb.ingress.kubernetes.io/target-type: ip
alb.ingress.kubernetes.io/healthcheck-path: /health
spec:
ingressClassName: alb
rules:
- host: myapp.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: myapp-stable
port:
number: 80
Traefik
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-traefik
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
ports:
- containerPort: 8080
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
traefik:
# Traefik weighted round robin
weightedTraefikServiceName: myapp-weighted
steps:
- setWeight: 20
- pause: {duration: 5m}
- setWeight: 50
- pause: {duration: 10m}
- setWeight: 80
- pause: {duration: 10m}
---
# Traefik IngressRoute
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: myapp-route
spec:
entryPoints:
- web
routes:
- match: Host(`myapp.example.com`)
kind: Rule
services:
- name: myapp-weighted
port: 80
---
# Traefik weighted service (managed by Rollout)
apiVersion: traefik.io/v1alpha1
kind: TraefikService
metadata:
name: myapp-weighted
spec:
weighted:
services:
- name: myapp-stable
port: 80
weight: 100
- name: myapp-canary
port: 80
weight: 0
Header-Based Routing
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: header-routing
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
nginx:
stableIngress: myapp-ingress
# Enable header-based routing
additionalIngressAnnotations:
# Users with X-Canary: always header get canary
canary-by-header: X-Canary
canary-by-header-value: "always"
steps:
# Start with header-only routing (0% weight)
- setWeight: 0
- pause: {duration: 10m} # Test with header
# Begin percentage-based rollout
- setWeight: 10
- pause: {duration: 5m}
- setWeight: 30
- pause: {duration: 5m}
- setWeight: 50
- pause: {}
# Test with: curl -H "X-Canary: always" https://myapp.example.com
Integration with Service Meshes
Integrating Argo Rollouts with service meshes for advanced traffic management.
Key Concepts
- Virtual Service: Service mesh routing rules
- Destination Rule: Service mesh subset definitions
- Traffic Splitting: Mesh-level traffic control
- Mutual TLS: Secure service-to-service communication
flowchart TD
subgraph "Service Mesh Integration"
A[Client] --> B[Virtual Service]
B --> C{Traffic Split}
C -->|80%| D[Stable Subset]
C -->|20%| E[Canary Subset]
D --> F[Stable Pods v1]
E --> G[Canary Pods v2]
H[Rollout Controller] --> B
H --> I[Update Weights]
end
Istio Integration
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-istio
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
version: v2 # Version label for Istio
spec:
containers:
- name: myapp
image: myapp:v2.0.0
ports:
- containerPort: 8080
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
istio:
# Virtual service to manage
virtualService:
name: myapp-vsvc
# Optional: Specific routes to update
routes:
- primary
# Optional: Destination rule with subsets
destinationRule:
name: myapp-destrule
canarySubsetName: canary
stableSubsetName: stable
steps:
- setWeight: 10
- pause: {duration: 5m}
- setWeight: 25
- pause: {duration: 5m}
- setWeight: 50
- pause: {duration: 10m}
- setWeight: 75
- pause: {duration: 10m}
---
# Istio Virtual Service
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: myapp-vsvc
spec:
hosts:
- myapp.example.com
gateways:
- myapp-gateway
http:
- name: primary
route:
- destination:
host: myapp-stable
subset: stable
weight: 100
- destination:
host: myapp-canary
subset: canary
weight: 0
---
# Istio Destination Rule
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: myapp-destrule
spec:
host: myapp-stable
subsets:
- name: stable
labels:
version: v1
- name: canary
labels:
version: v2
trafficPolicy:
tls:
mode: ISTIO_MUTUAL
---
# Gateway
apiVersion: networking.istio.io/v1
kind: Gateway
metadata:
name: myapp-gateway
spec:
selector:
istio: ingressgateway
servers:
- port:
number: 80
name: http
protocol: HTTP
hosts:
- myapp.example.com
Linkerd Integration (SMI)
Note: Linkerd 2.11+ no longer ships the SMI
TrafficSplitCRD. On modern Linkerd you must install the SMI extension/CRD separately (or use a different mesh) for thissmitraffic routing to work.
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp-linkerd
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
annotations:
# Linkerd proxy injection
linkerd.io/inject: enabled
spec:
containers:
- name: myapp
image: myapp:v2.0.0
ports:
- containerPort: 8080
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
smi:
# SMI TrafficSplit resource
trafficSplitName: myapp-trafficsplit
# Optional: Root service
rootService: myapp-root
steps:
- setWeight: 10
- pause: {duration: 5m}
- setWeight: 30
- pause: {duration: 10m}
- setWeight: 50
- pause: {}
---
# SMI Traffic Split (managed by Rollout)
apiVersion: split.smi-spec.io/v1alpha1
kind: TrafficSplit
metadata:
name: myapp-trafficsplit
spec:
service: myapp-root
backends:
- service: myapp-stable
weight: 100
- service: myapp-canary
weight: 0
---
# Root service
apiVersion: v1
kind: Service
metadata:
name: myapp-root
spec:
selector:
app: myapp
ports:
- port: 80
targetPort: 8080
Istio with Header-Based Routing
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: advanced-istio
spec:
replicas: 10
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
version: v2
spec:
containers:
- name: myapp
image: myapp:v2.0.0
strategy:
canary:
stableService: myapp-stable
canaryService: myapp-canary
trafficRouting:
istio:
virtualService:
name: myapp-vsvc
routes:
- primary
steps:
# Start with header-based routing
- setHeaderRoute:
name: test-users
match:
- headerName: X-Test-User
headerValue:
exact: "true"
- pause: {duration: 10m}
# Begin percentage rollout
- setWeight: 10
- pause: {duration: 5m}
- setWeight: 25
- pause: {duration: 5m}
- setWeight: 50
- pause: {}
---
# Istio Virtual Service with header matching
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: myapp-vsvc
spec:
hosts:
- myapp.example.com
http:
# Header-based route (managed by Rollout)
- name: test-users
match:
- headers:
X-Test-User:
exact: "true"
route:
- destination:
host: myapp-canary
weight: 100
# Primary route (managed by Rollout)
- name: primary
route:
- destination:
host: myapp-stable
weight: 100
- destination:
host: myapp-canary
weight: 0
Quick Reference
Common Commands
| Command | Description |
|---|---|
kubectl argo rollouts list |
List all rollouts |
kubectl argo rollouts get rollout <name> |
Get rollout status |
kubectl argo rollouts get rollout <name> --watch |
Watch rollout progress |
kubectl argo rollouts promote <name> |
Promote to next step |
kubectl argo rollouts promote <name> --full |
Skip to end |
kubectl argo rollouts pause <name> |
Pause rollout |
kubectl argo rollouts resume <name> |
Resume paused rollout |
kubectl argo rollouts abort <name> |
Abort rollout |
kubectl argo rollouts retry rollout <name> |
Retry failed rollout |
kubectl argo rollouts undo <name> |
Rollback to previous |
kubectl argo rollouts undo <name> --to-revision=N |
Rollback to revision N |
kubectl argo rollouts set image <name> <container>=<image> |
Update image |
kubectl argo rollouts restart <name> |
Restart rollout |
kubectl argo rollouts status <name> |
Get rollout status |
kubectl argo rollouts dashboard |
Launch dashboard |
Rollout Status Values
| Status | Description |
|---|---|
| Healthy | All replicas healthy and at desired version |
| Progressing | Rollout in progress |
| Degraded | Rollout has failed or exceeded deadline |
| Paused | Rollout paused waiting for promotion |
Analysis Result Values
| Result | Description |
|---|---|
| Successful | Analysis passed all metrics |
| Failed | Analysis failed metrics threshold |
| Error | Analysis encountered an error |
| Running | Analysis currently executing |
| Pending | Analysis not yet started |
| Inconclusive | Cannot determine result |
Metrics Providers
| Provider | Use Case |
|---|---|
| Prometheus | Open-source metrics and alerting |
| Datadog | Cloud monitoring and analytics |
| New Relic | Application performance monitoring |
| Wavefront | Cloud-native monitoring |
| Kayenta | Statistical analysis engine |
| Web | HTTP endpoint checks |
| Job | Kubernetes job-based tests |
| CloudWatch | AWS metrics |
| Graphite | Time-series database |
Traffic Routing Options
| Method | Description |
|---|---|
| Nginx Ingress | NGINX-based traffic splitting |
| ALB | AWS Application Load Balancer |
| Istio | Service mesh with virtual services |
| Linkerd | Lightweight service mesh (SMI) |
| Traefik | Cloud-native edge router |
| App Mesh | AWS service mesh |
| Ambassador | API Gateway with traffic control |
Step Types Reference
| Step Type | Description | Example |
|---|---|---|
setWeight |
Set traffic weight percentage | setWeight: 20 |
pause |
Pause with duration or indefinitely | pause: {duration: 5m} or pause: {} |
setCanaryScale |
Set canary replica count | setCanaryScale: {replicas: 3} |
analysis |
Run inline analysis | analysis: {templates: [...]} |
experiment |
Run A/B test experiment | experiment: {duration: 10m} |
setHeaderRoute |
Route by header (Istio) | setHeaderRoute: {name: test} |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Rollout stuck in Progressing | Analysis timeout or failure | Check AnalysisRun: kubectl get analysisrun -n <namespace> |
| Traffic not splitting | Ingress/mesh misconfiguration | Verify service names and traffic routing config match |
| Analysis always fails | Incorrect query or threshold | Test query directly against metrics provider |
| Pods not scaling | Resource constraints | Check cluster capacity: kubectl describe nodes |
| Rollout not starting | Deployment hasn't changed | Update pod template to trigger new revision |
| Cannot promote | No manual pause step | Ensure strategy includes pause: {} step |
| Service mesh integration fails | Missing VirtualService/DestinationRule | Create required mesh resources first |
| Analysis provider timeout | Network or authentication issues | Check provider connectivity and credentials |
| Rollback not working | No previous revision | Ensure revisionHistoryLimit > 1 |
| Dashboard not showing rollout | Namespace filtering | Check dashboard is monitoring correct namespace |
| Canary service has no endpoints | Selector doesn't match pods | Verify service selector matches pod labels |
| Analysis inconclusive | No data from provider | Check metric exists and query returns data |
Debugging Commands
# View rollout events
kubectl describe rollout <name> -n <namespace>
# Check analysis runs
kubectl get analysisrun -n <namespace>
kubectl describe analysisrun <name> -n <namespace>
# View controller logs
kubectl logs -n argo-rollouts deployment/argo-rollouts -f
# Check ReplicaSet status
kubectl get rs -l app=<app-name> -n <namespace>
# View pod status and labels
kubectl get pods -l app=<app-name> -n <namespace> --show-labels
# Check traffic routing config
kubectl get ingress <name> -n <namespace> -o yaml
kubectl get virtualservice <name> -n <namespace> -o yaml
# Export rollout for inspection
kubectl get rollout <name> -n <namespace> -o yaml > rollout-debug.yaml
# Validate analysis template
kubectl apply --dry-run=client -f analysis-template.yaml
# Check service endpoints
kubectl get endpoints <service-name> -n <namespace>
# View rollout in JSON format
kubectl argo rollouts get rollout <name> -n <namespace> -o json
Troubleshooting Analysis
# List all analysis runs
kubectl get analysisrun -n <namespace>
# Get analysis run details
kubectl describe analysisrun <name> -n <namespace>
# View analysis run status
kubectl get analysisrun <name> -n <namespace> -o jsonpath='{.status}'
# Check metric results
kubectl get analysisrun <name> -n <namespace> -o jsonpath='{.status.metricResults}'
# Test Prometheus query manually
kubectl run -it --rm debug --image=curlimages/curl --restart=Never -- \
curl "http://prometheus:9090/api/v1/query?query=<encoded-query>"
# Delete failed analysis run to retry
kubectl delete analysisrun <name> -n <namespace>
# Watch analysis run progress
kubectl get analysisrun <name> -n <namespace> -w
# Check analysis template args
kubectl get analysisrun <name> -n <namespace> -o jsonpath='{.spec.args}'
Troubleshooting Traffic Routing
# NGINX Ingress: Check canary ingress creation
kubectl get ingress -n <namespace> | grep canary
# NGINX Ingress: View canary annotations
kubectl get ingress <name>-<hash>-canary -n <namespace> -o yaml
# Istio: Check VirtualService weights
kubectl get virtualservice <name> -n <namespace> -o yaml | grep -A 5 weight
# Linkerd: Check TrafficSplit status
kubectl get trafficsplit <name> -n <namespace> -o yaml
# Check service selectors
kubectl get svc <name> -n <namespace> -o yaml | grep -A 5 selector
# Verify pod labels match service
kubectl get pods -l app=<name> -n <namespace> --show-labels
Related Topics
The following topics complement Argo Rollouts and are commonly used together in progressive delivery workflows:
- ArgoCD - GitOps continuous delivery that integrates seamlessly with Argo Rollouts for deployment automation
- Istio - Service mesh providing advanced traffic management capabilities for canary deployments
- Prometheus - Metrics collection and querying for automated analysis during rollouts
- Kubernetes - Understanding Kubernetes services, deployments, and networking is essential for Rollouts
- Grafana - Visualising rollout metrics and creating dashboards for deployment monitoring
- NGINX Ingress - Popular ingress controller for traffic splitting during canary deployments