Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Argo Rollouts

Progressive delivery controller for Kubernetes providing advanced deployment strategies like canary and blue/green.

Argo Rollouts

Progressive delivery controller for Kubernetes providing advanced deployment strategies like canary and blue/green.

Overview

Argo Rollouts is a Kubernetes controller that provides advanced deployment capabilities beyond standard Kubernetes Deployments. It enables progressive delivery through canary deployments, blue/green deployments, and automated analysis to minimise risk during application updates. Rollouts integrates with ingress controllers and service meshes to gradually shift traffic, and supports automated rollback based on metrics from Prometheus, Datadog, New Relic, and other providers.

Progressive Delivery WorkflowCanaryBlue/GreenYesNoNew VersionRollout ControllerStrategy?Gradual TrafficShiftFull SwitchAnalysisMetrics OK?Continue/CompleteAutomatic Rollback100% New Version100% Old VersionProgressive Delivery WorkflowCanaryBlue/GreenYesNoNew VersionRollout ControllerStrategy?Gradual TrafficShiftFull SwitchAnalysisMetrics OK?Continue/CompleteAutomatic Rollback100% New Version100% Old Version

Installation and Setup

Argo Rollouts consists of a controller that watches Rollout resources and manages progressive delivery.

Key Concepts

  • Controller: Manages Rollout resources and orchestrates deployments
  • Kubectl Plugin: CLI tool for managing and monitoring rollouts
  • Dashboard: Optional UI for visualising rollout status
  • CRDs: Custom Resource Definitions for Rollout, AnalysisTemplate, AnalysisRun, Experiment
Argo Rollouts ArchitectureRollout ControllerRollout CRDAnalysisRunReplicaSetsStrategy ConfigTraffic ManagementMetrics ProvidersIngress/Service MeshPodsArgo Rollouts ArchitectureRollout ControllerRollout CRDAnalysisRunReplicaSetsStrategy ConfigTraffic ManagementMetrics ProvidersIngress/Service MeshPods

Installation Commands

# Install Argo Rollouts controller
kubectl create namespace argo-rollouts
kubectl apply -n argo-rollouts -f https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml

# Install via Helm
helm repo add argo https://argoproj.github.io/argo-helm
helm install argo-rollouts argo/argo-rollouts -n argo-rollouts --create-namespace

# Install kubectl plugin (Linux)
curl -LO https://github.com/argoproj/argo-rollouts/releases/latest/download/kubectl-argo-rollouts-linux-amd64
chmod +x kubectl-argo-rollouts-linux-amd64
sudo mv kubectl-argo-rollouts-linux-amd64 /usr/local/bin/kubectl-argo-rollouts

# Install kubectl plugin (macOS)
brew install argoproj/tap/kubectl-argo-rollouts

# Verify installation
kubectl argo rollouts version

# Install dashboard (optional)
kubectl apply -n argo-rollouts -f https://github.com/argoproj/argo-rollouts/releases/latest/download/dashboard-install.yaml

# Access dashboard
kubectl port-forward -n argo-rollouts svc/argo-rollouts-dashboard 3100:3100
# Visit http://localhost:3100

Controller Configuration

apiVersion: apps/v1
kind: Deployment
metadata:
  name: argo-rollouts
  namespace: argo-rollouts
spec:
  replicas: 2
  selector:
    matchLabels:
      app.kubernetes.io/name: argo-rollouts
  template:
    metadata:
      labels:
        app.kubernetes.io/name: argo-rollouts
    spec:
      serviceAccountName: argo-rollouts
      containers:
        - name: argo-rollouts
          image: quay.io/argoproj/argo-rollouts:stable
          args:
            - --loglevel
            - info
            - --leader-elect
            - --namespaced  # Watch specific namespaces
          ports:
            - containerPort: 8090
              name: metrics
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
            limits:
              cpu: 500m
              memory: 512Mi
          securityContext:
            allowPrivilegeEscalation: false
            runAsNonRoot: true
            capabilities:
              drop:
                - ALL

Rollout Objects and Strategies

Rollouts are Kubernetes custom resources that extend Deployment functionality with progressive delivery strategies.

Key Concepts

  • Rollout: Similar to Deployment but with advanced deployment strategies
  • Strategy: Defines how updates are rolled out (Canary, Blue/Green)
  • Revision: Each update creates a new revision tracked by the controller
  • Stable/Canary ReplicaSets: Managed automatically during progressive delivery
  • Pod Template Hash: Label used to track versions
Rollout StructureRolloutStrategyPod TemplateReplicasCanary StepsBlue/Green ConfigAnalysisRollout StructureRolloutStrategyPod TemplateReplicasCanary StepsBlue/Green ConfigAnalysis

Basic Rollout Manifest

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp
  namespace: production
spec:
  replicas: 5

  # Revision history to keep
  revisionHistoryLimit: 3

  # Selector must match pod template labels
  selector:
    matchLabels:
      app: myapp

  # Pod template (same as Deployment)
  template:
    metadata:
      labels:
        app: myapp
        version: stable
    spec:
      containers:
        - name: myapp
          image: myapp:v1.0.0
          ports:
            - containerPort: 8080
              name: http
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
            limits:
              cpu: 500m
              memory: 512Mi
          livenessProbe:
            httpGet:
              path: /health
              port: 8080
            initialDelaySeconds: 30
            periodSeconds: 10
          readinessProbe:
            httpGet:
              path: /ready
              port: 8080
            initialDelaySeconds: 10
            periodSeconds: 5

  # Strategy defined separately (see following sections)
  strategy:
    canary: {}

Common Commands

# List rollouts
kubectl argo rollouts list -n production

# Get rollout status
kubectl argo rollouts get rollout myapp -n production

# Watch rollout progress
kubectl argo rollouts get rollout myapp -n production --watch

# Describe rollout
kubectl describe rollout myapp -n production

# View rollout history
kubectl argo rollouts history rollout myapp -n production

# Trigger a rollout by updating image
kubectl argo rollouts set image myapp myapp=myapp:v2.0.0 -n production

# Manually promote a rollout
kubectl argo rollouts promote myapp -n production

# Skip all remaining steps and deploy immediately
kubectl argo rollouts promote myapp -n production --full

# Pause a rollout
kubectl argo rollouts pause myapp -n production

# Resume a paused rollout
kubectl argo rollouts resume myapp -n production

# Abort a rollout
kubectl argo rollouts abort myapp -n production

# Retry a failed rollout
kubectl argo rollouts retry rollout myapp -n production

# Undo to previous revision
kubectl argo rollouts undo myapp -n production

# Undo to specific revision
kubectl argo rollouts undo myapp --to-revision=3 -n production

# Restart rollout (create new revision)
kubectl argo rollouts restart myapp -n production

# Get rollout status (non-interactive)
kubectl argo rollouts status myapp -n production

Converting Deployment to Rollout

# Before (Deployment)
apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  replicas: 5
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v1.0.0

---
# After (Rollout)
apiVersion: argoproj.io/v1alpha1
kind: Rollout  # Changed from Deployment
metadata:
  name: myapp
spec:
  replicas: 5
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v1.0.0
  strategy:  # Added strategy
    canary:
      steps:
        - setWeight: 20
        - pause: {duration: 1m}
        - setWeight: 50
        - pause: {duration: 2m}

Canary Deployment Strategy

Canary deployments gradually shift traffic to the new version in controlled steps, allowing validation at each stage.

Key Concepts

  • Traffic Weight: Percentage of traffic sent to canary version
  • Steps: Sequence of weight changes and pauses
  • Manual/Auto Promotion: Whether steps advance automatically or require approval
  • MaxSurge/MaxUnavailable: Control pod creation during rollout
  • Canary Service: Service pointing to canary pods for testing
AnalysisCanary ReplicaSetStable ReplicaSetRollout ControllerUserAnalysisCanary ReplicaSetStable ReplicaSetRollout ControllerUserManual gate - wait for promotionUpdate imageCreate v2 (10% weight)Start metrics checkMetrics passScale to 25% weightContinue monitoringMetrics passScale to 50% weightPromoteScale to 100% weightScale down v1Rollout completeAnalysisCanary ReplicaSetStable ReplicaSetRollout ControllerUserAnalysisCanary ReplicaSetStable ReplicaSetRollout ControllerUserManual gate - wait for promotionUpdate imageCreate v2 (10% weight)Start metrics checkMetrics passScale to 25% weightContinue monitoringMetrics passScale to 50% weightPromoteScale to 100% weightScale down v1Rollout complete

Canary Strategy Configuration

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-canary
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v1.0.0
          ports:
            - containerPort: 8080

  strategy:
    canary:
      # Maximum additional replicas during rollout
      maxSurge: "25%"
      # Maximum unavailable replicas during rollout
      maxUnavailable: 0

      # Canary steps
      steps:
        # Step 1: Send 10% of traffic to canary
        - setWeight: 10
        - pause: {duration: 5m}

        # Step 2: Increase to 25%
        - setWeight: 25
        - pause: {duration: 5m}

        # Step 3: Increase to 50%
        - setWeight: 50
        - pause: {duration: 10m}

        # Step 4: Increase to 75%
        - setWeight: 75
        - pause: {duration: 10m}

        # Step 5: Manual promotion pause (wait indefinitely)
        - pause: {}

        # Final: 100% to new version (implicit)

      # Optional: Analysis at specific steps
      analysis:
        templates:
          - templateName: success-rate
        startingStep: 2  # Start analysis at step index 2

      # Traffic routing (see integration section)
      trafficRouting:
        nginx:
          stableIngress: myapp-stable
          annotationPrefix: nginx.ingress.kubernetes.io

Canary with Background Analysis

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-auto-analysis
spec:
  replicas: 5
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v1.0.0

  strategy:
    canary:
      steps:
        - setWeight: 20
        - pause: {duration: 2m}
        - setWeight: 40
        - pause: {duration: 2m}
        - setWeight: 60
        - pause: {duration: 2m}
        - setWeight: 80
        - pause: {duration: 2m}

      # Background analysis runs throughout the rollout
      analysis:
        templates:
          - templateName: error-rate
          - templateName: response-time
        # Analysis runs for entire canary deployment
        args:
          - name: service-name
            value: myapp
          - name: canary-hash
            valueFrom:
              podTemplateHashValue: Latest

Canary with Manual Gates

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: production-app
spec:
  replicas: 10
  selector:
    matchLabels:
      app: production-app
  template:
    metadata:
      labels:
        app: production-app
    spec:
      containers:
        - name: app
          image: production-app:v2.0.0

  strategy:
    canary:
      maxSurge: "20%"
      maxUnavailable: 0

      steps:
        # Initial canary with 5%
        - setWeight: 5
        - pause: {duration: 5m}

        # Expand to 10% with manual approval
        - setWeight: 10
        - pause: {}  # Wait indefinitely for promotion

        # Progressive rollout with analysis
        - setWeight: 25
        - pause: {duration: 10m}

        - setWeight: 50
        - pause: {}  # Another manual gate

        - setWeight: 75
        - pause: {duration: 10m}

        # Final manual approval before 100%
        - pause: {}

      analysis:
        templates:
          - templateName: critical-metrics
        startingStep: 2

Blue/Green Deployment Strategy

Blue/green deployments maintain two complete environments, switching traffic atomically between them.

Key Concepts

  • Active Service: Service pointing to stable (blue) version
  • Preview Service: Service pointing to new (green) version for testing
  • Automatic Promotion: Can automatically promote after duration
  • Instant Switchover: Traffic switches immediately, no gradual shift
  • Preview Testing: Test new version before promoting to active
Blue/Green Deployment FlowPassFailPassFailCurrent: Blue v1Deploy Green v2Preview Servicepoints to GreenRun Pre-PromotionAnalysisSwitch ActiveServiceto GreenAbort - Stay on BlueRun Post-PromotionAnalysisScale Down BlueRollback to BlueCompleteBlue Still ActiveBlue/Green Deployment FlowPassFailPassFailCurrent: Blue v1Deploy Green v2Preview Servicepoints to GreenRun Pre-PromotionAnalysisSwitch ActiveServiceto GreenAbort - Stay on BlueRun Post-PromotionAnalysisScale Down BlueRollback to BlueCompleteBlue Still Active

Blue/Green Strategy Configuration

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-bluegreen
spec:
  replicas: 5

  selector:
    matchLabels:
      app: myapp

  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v1.0.0
          ports:
            - containerPort: 8080

  strategy:
    blueGreen:
      # Active service (production traffic)
      activeService: myapp-active

      # Preview service (for testing new version)
      previewService: myapp-preview

      # Automatically promote after duration
      autoPromotionEnabled: false
      autoPromotionSeconds: 300  # 5 minutes

      # Scale down old version after promotion
      scaleDownDelaySeconds: 30
      scaleDownDelayRevisionLimit: 2

      # Anti-affinity between blue and green
      antiAffinity:
        requiredDuringSchedulingIgnoredDuringExecution: {}

      # Optional: Run analysis before promotion
      prePromotionAnalysis:
        templates:
          - templateName: smoke-tests
        args:
          - name: service-url
            value: http://myapp-preview

      # Optional: Run analysis after promotion
      postPromotionAnalysis:
        templates:
          - templateName: integration-tests
        args:
          - name: service-url
            value: http://myapp-active

Required Services for Blue/Green

# Active service - receives production traffic
apiVersion: v1
kind: Service
metadata:
  name: myapp-active
spec:
  selector:
    app: myapp
    # Rollout controller manages this label automatically
  ports:
    - port: 80
      targetPort: 8080
      protocol: TCP
  type: ClusterIP

---
# Preview service - for testing new version
apiVersion: v1
kind: Service
metadata:
  name: myapp-preview
spec:
  selector:
    app: myapp
    # Rollout controller manages this label automatically
  ports:
    - port: 80
      targetPort: 8080
      protocol: TCP
  type: ClusterIP

Blue/Green with Comprehensive Analysis

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: critical-app
spec:
  replicas: 10

  selector:
    matchLabels:
      app: critical-app

  template:
    metadata:
      labels:
        app: critical-app
    spec:
      containers:
        - name: app
          image: critical-app:v2.0.0
          ports:
            - containerPort: 8080

  strategy:
    blueGreen:
      activeService: critical-app-active
      previewService: critical-app-preview

      # Manual promotion only
      autoPromotionEnabled: false

      # Keep old version for 1 hour after promotion
      scaleDownDelaySeconds: 3600
      scaleDownDelayRevisionLimit: 1

      # Pre-promotion smoke tests
      prePromotionAnalysis:
        templates:
          - templateName: smoke-tests
          - templateName: load-tests
        args:
          - name: preview-url
            value: http://critical-app-preview
          - name: duration
            value: "300"

      # Post-promotion verification
      postPromotionAnalysis:
        templates:
          - templateName: error-rate-check
          - templateName: performance-check
        args:
          - name: active-url
            value: http://critical-app-active
          - name: threshold
            value: "0.01"  # 1% error rate

Canary vs Blue/Green Comparison

Understanding when to use each strategy based on your requirements.

Strategy Comparison

Aspect Canary Blue/Green
Traffic Shift Gradual (percentage-based) Instant (atomic switch)
Resource Usage Efficient (gradual scaling) Higher (2x pods during switch)
Rollback Speed Immediate (shift back) Immediate (switch back)
Testing Ability Production traffic sampling Full preview environment
Complexity Higher (traffic management) Lower (simple switch)
Risk Lower (gradual exposure) Higher (instant full exposure)
Best For Production deployments Staging, critical systems
User Impact Gradual Immediate
Infrastructure Needs ingress/mesh Standard services
Validation Real user traffic Isolated testing
Granularity Fine-grained control All-or-nothing
Blue/Green Deployment100% Blue v1Test Green v2via PreviewSwitch: 100% Greenv2Scale down BlueCanary Deployment100% v190% v1, 10% v275% v1, 25% v250% v1, 50% v225% v1, 75% v2100% v2Blue/Green Deployment100% Blue v1Test Green v2via PreviewSwitch: 100% Greenv2Scale down BlueCanary Deployment100% v190% v1, 10% v275% v1, 25% v250% v1, 50% v225% v1, 75% v2100% v2

Decision Matrix

YesNoLow risk toleranceYesNoHigh risk toleranceYesNoChoose DeploymentStrategyNeed previewenvironment?Blue/GreenRisk tolerance?Have ingress/service mesh?CanaryNeed gradualrollout?Canary StrategyBlue/Green StrategyYesNoLow risk toleranceYesNoHigh risk toleranceYesNoChoose DeploymentStrategyNeed previewenvironment?Blue/GreenRisk tolerance?Have ingress/service mesh?CanaryNeed gradualrollout?Canary StrategyBlue/Green Strategy

Use Case Recommendations

Use Canary When:

  • Deploying to production with real user traffic
  • Want to limit blast radius of potential issues
  • Have metrics to validate new version
  • Can integrate with ingress controller or service mesh
  • Need fine-grained control over traffic shift
  • Cost-conscious (gradual scaling more efficient)
  • Want to validate with real traffic patterns
  • Need to detect issues early with minimal impact

Use Blue/Green When:

  • Need isolated preview environment for testing
  • Require instant cutover (e.g., coordinated releases)
  • Don't have advanced traffic routing available
  • Deploy to staging/pre-production environments
  • Need simple, deterministic deployment process
  • Can afford temporary resource doubling
  • Require quick, complete rollback capability
  • Want to run comprehensive tests before cutover
  • Database schema changes require atomic switch

Hybrid Approach

# Staging: Blue/Green for quick testing
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-staging
  namespace: staging
spec:
  replicas: 3
  selector:
    matchLabels:
      app: myapp
      env: staging
  template:
    metadata:
      labels:
        app: myapp
        env: staging
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    blueGreen:
      activeService: myapp-staging
      previewService: myapp-staging-preview
      autoPromotionEnabled: true
      autoPromotionSeconds: 60
      prePromotionAnalysis:
        templates:
          - templateName: basic-checks

---
# Production: Canary for careful rollout
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-production
  namespace: production
spec:
  replicas: 20
  selector:
    matchLabels:
      app: myapp
      env: production
  template:
    metadata:
      labels:
        app: myapp
        env: production
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      maxSurge: "25%"
      maxUnavailable: 0
      steps:
        - setWeight: 5
        - pause: {duration: 10m}
        - setWeight: 20
        - pause: {duration: 10m}
        - setWeight: 50
        - pause: {}  # Manual gate
        - setWeight: 80
        - pause: {duration: 15m}
      analysis:
        templates:
          - templateName: error-rate
          - templateName: latency
        startingStep: 1

Analysis Templates and Metrics Providers

Analysis templates define success criteria for deployments using metrics from various providers.

Key Concepts

  • AnalysisTemplate: Reusable metric query and success criteria definition
  • AnalysisRun: Instance of analysis execution during rollout
  • Metrics Provider: Backend system providing metrics (Prometheus, Datadog, etc.)
  • Success Criteria: Thresholds determining pass/fail
  • Arguments: Parameters passed from Rollout to template
Analysis WorkflowSuccessFailureErrorRollout TriggersCreate AnalysisRunAnalysisTemplateDefine MetricsQuery ProviderEvaluateCriteriaContinue RolloutAbort & RollbackInconclusiveRetry or FailAnalysis WorkflowSuccessFailureErrorRollout TriggersCreate AnalysisRunAnalysisTemplateDefine MetricsQuery ProviderEvaluateCriteriaContinue RolloutAbort & RollbackInconclusiveRetry or Fail

AnalysisTemplate Structure

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: success-rate
  namespace: production
spec:
  # Arguments passed from Rollout
  args:
    - name: service-name
    - name: namespace
      value: production  # Default value

  # Metrics to evaluate
  metrics:
    - name: success-rate
      # How long to run analysis
      interval: 1m
      # Number of measurements
      count: 5
      # Maximum consecutive failures before abort
      failureLimit: 2
      # Initial delay before starting
      initialDelay: 30s

      # Success criteria
      successCondition: result >= 0.95
      failureCondition: result < 0.90

      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{
              service="{{args.service-name}}",
              namespace="{{args.namespace}}",
              status=~"2.."
            }[5m]))
            /
            sum(rate(http_requests_total{
              service="{{args.service-name}}",
              namespace="{{args.namespace}}"
            }[5m]))

Prometheus Provider

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: prometheus-metrics
spec:
  args:
    - name: service
    - name: threshold
      value: "0.99"

  metrics:
    # Error rate check
    - name: error-rate
      interval: 30s
      count: 10
      successCondition: result < 0.01  # Less than 1% error rate
      failureCondition: result >= 0.05  # More than 5% error rate
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{
              service="{{args.service}}",
              status=~"5.."
            }[2m]))
            /
            sum(rate(http_requests_total{
              service="{{args.service}}"
            }[2m]))

    # Response time check (p95)
    - name: response-time-p95
      interval: 30s
      count: 10
      successCondition: result < 500  # Less than 500ms
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            histogram_quantile(0.95,
              sum(rate(http_request_duration_seconds_bucket{
                service="{{args.service}}"
              }[2m])) by (le)
            ) * 1000

    # CPU usage check
    - name: cpu-usage
      interval: 1m
      count: 5
      successCondition: result < 80  # Less than 80%
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(container_cpu_usage_seconds_total{
              pod=~"{{args.service}}.*"
            }[2m])) by (pod)
            /
            sum(container_spec_cpu_quota{
              pod=~"{{args.service}}.*"
            } / container_spec_cpu_period{
              pod=~"{{args.service}}.*"
            }) by (pod)
            * 100

    # Memory usage check
    - name: memory-usage
      interval: 1m
      count: 5
      successCondition: result < 85  # Less than 85%
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(container_memory_working_set_bytes{
              pod=~"{{args.service}}.*"
            }) by (pod)
            /
            sum(container_spec_memory_limit_bytes{
              pod=~"{{args.service}}.*"
            }) by (pod)
            * 100

Datadog Provider

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: datadog-metrics
spec:
  args:
    - name: service-name

  metrics:
    - name: error-rate
      interval: 1m
      count: 5
      successCondition: result < 5
      provider:
        datadog:
          apiVersion: v1
          interval: 5m
          query: |
            sum:trace.http.request.errors{
              service:{{args.service-name}}
            }.as_count()
            /
            sum:trace.http.request.hits{
              service:{{args.service-name}}
            }.as_count()
            * 100

    - name: avg-response-time
      interval: 1m
      count: 5
      successCondition: result < 0.5  # 500ms
      provider:
        datadog:
          apiVersion: v1
          interval: 5m
          query: |
            avg:trace.http.request.duration{
              service:{{args.service-name}}
            }

    - name: apm-errors
      interval: 2m
      count: 5
      successCondition: result < 10
      provider:
        datadog:
          apiVersion: v1
          interval: 5m
          query: |
            sum:trace.errors{
              service:{{args.service-name}}
            }.as_count()

New Relic Provider

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: newrelic-metrics
spec:
  args:
    - name: service-name
    - name: canary-hash

  metrics:
    - name: apdex-score
      interval: 2m
      count: 5
      successCondition: result > 0.95
      provider:
        newRelic:
          profile: production-account
          query: |
            FROM Transaction
            SELECT apdex(duration, t: 0.5)
            WHERE appName = '{{args.service-name}}'
            AND tags.revision = '{{args.canary-hash}}'

    - name: error-percentage
      interval: 2m
      count: 5
      successCondition: result < 1
      provider:
        newRelic:
          profile: production-account
          query: |
            FROM Transaction
            SELECT percentage(count(*), WHERE error IS true)
            WHERE appName = '{{args.service-name}}'
            AND tags.revision = '{{args.canary-hash}}'

    - name: throughput
      interval: 2m
      count: 5
      successCondition: result > 100
      provider:
        newRelic:
          profile: production-account
          query: |
            FROM Transaction
            SELECT rate(count(*), 1 minute)
            WHERE appName = '{{args.service-name}}'
            AND tags.revision = '{{args.canary-hash}}'

Web/HTTP Provider

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: http-health-check
spec:
  args:
    - name: service-url

  metrics:
    - name: health-check
      interval: 30s
      count: 10
      successCondition: result == "200"
      failureLimit: 3
      provider:
        web:
          url: "{{args.service-url}}/health"
          headers:
            - key: User-Agent
              value: argo-rollouts
          jsonPath: "{$.status}"  # Parse JSON response
          timeoutSeconds: 10

    - name: api-response-check
      interval: 1m
      count: 5
      successCondition: "result.healthy == true"
      provider:
        web:
          url: "{{args.service-url}}/api/status"
          jsonPath: "{$}"

    - name: metrics-endpoint
      interval: 1m
      count: 5
      successCondition: result.error_rate < 0.01
      provider:
        web:
          url: "{{args.service-url}}/metrics"
          jsonPath: "{$.error_rate}"

Job Provider

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: load-test
spec:
  args:
    - name: target-url

  metrics:
    - name: load-test-results
      provider:
        job:
          spec:
            backoffLimit: 0
            template:
              spec:
                containers:
                  - name: load-test
                    image: grafana/k6:latest
                    command:
                      - k6
                      - run
                      - /scripts/load-test.js
                    env:
                      - name: TARGET_URL
                        value: "{{args.target-url}}"
                    volumeMounts:
                      - name: script
                        mountPath: /scripts
                volumes:
                  - name: script
                    configMap:
                      name: k6-load-test-script
                restartPolicy: Never

---
apiVersion: v1
kind: ConfigMap
metadata:
  name: k6-load-test-script
data:
  load-test.js: |
    import http from 'k6/http';
    import { check, sleep } from 'k6';

    export const options = {
      stages: [
        { duration: '1m', target: 50 },
        { duration: '2m', target: 50 },
        { duration: '1m', target: 0 },
      ],
      thresholds: {
        http_req_duration: ['p(95)<500'],
        http_req_failed: ['rate<0.01'],
      },
    };

    export default function () {
      const res = http.get(__ENV.TARGET_URL);
      check(res, {
        'status is 200': (r) => r.status === 200,
      });
      sleep(1);
    }

Combining Multiple Templates

# Rollout using multiple analysis templates
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: comprehensive-analysis
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      steps:
        - setWeight: 20
        - pause: {duration: 5m}
        - setWeight: 50
        - pause: {duration: 5m}

      analysis:
        templates:
          # Combine multiple templates
          - templateName: prometheus-metrics
          - templateName: datadog-metrics
          - templateName: http-health-check

        args:
          - name: service
            value: myapp
          - name: service-name
            value: myapp
          - name: service-url
            value: http://myapp-canary

Multi-Provider Analysis

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: comprehensive-checks
spec:
  args:
    - name: service
    - name: canary-hash

  metrics:
    # Prometheus: Error rate
    - name: error-rate
      interval: 2m
      count: 5
      successCondition: result < 0.01
      failureLimit: 2
      provider:
        prometheus:
          address: http://prometheus:9090
          query: |
            sum(rate(http_requests_total{
              service="{{args.service}}",
              rollouts_pod_template_hash="{{args.canary-hash}}",
              status=~"5.."
            }[5m]))
            /
            sum(rate(http_requests_total{
              service="{{args.service}}",
              rollouts_pod_template_hash="{{args.canary-hash}}"
            }[5m]))

    # New Relic: Apdex score
    - name: apdex
      interval: 2m
      count: 5
      successCondition: result > 0.95
      provider:
        newRelic:
          profile: my-account
          query: |
            FROM Transaction
            SELECT apdex(duration, t: 0.5)
            WHERE appName = '{{args.service}}'
            AND tags.revision = '{{args.canary-hash}}'

    # Datadog: Log errors
    - name: log-errors
      interval: 2m
      count: 5
      successCondition: result < 10
      provider:
        datadog:
          interval: 5m
          query: |
            sum:log.error.count{
              service:{{args.service}},
              revision:{{args.canary-hash}}
            }.as_count()

Progressive Delivery Steps and Pauses

Controlling the pace and gates of rollout progression through steps and pauses.

Key Concepts

  • Steps: Ordered sequence of actions during canary rollout
  • setWeight: Traffic percentage to send to canary
  • pause: Wait period (duration or indefinite)
  • experiment: A/B testing with multiple variants
  • analysis: Metric evaluation at specific steps
  • setCanaryScale: Control number of canary replicas independently
Start RolloutsetWeight 10%Duration elapsedMetrics passMetrics failsetWeight 50%Duration elapsedManual promoteManual abortsetWeight 100%Return to stableStep1Pause1Analysis1Step2RollbackPause2ManualGateStep3CompleteStart RolloutsetWeight 10%Duration elapsedMetrics passMetrics failsetWeight 50%Duration elapsedManual promoteManual abortsetWeight 100%Return to stableStep1Pause1Analysis1Step2RollbackPause2ManualGateStep3Complete

Step Types

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: step-examples
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      steps:
        # Set traffic weight
        - setWeight: 10

        # Pause for specific duration
        - pause:
            duration: 5m

        # Pause indefinitely (manual promotion required)
        - pause: {}

        # Set traffic weight with canary replicas
        - setWeight: 30
        - setCanaryScale:
            weight: 30  # Scale canary to 30% of spec.replicas

        # Or set explicit replica count
        - setCanaryScale:
            replicas: 3  # Exactly 3 canary replicas

        # Inline analysis
        - analysis:
            templates:
              - templateName: success-rate
            args:
              - name: service-name
                value: myapp

        # Continue progression
        - setWeight: 50
        - pause:
            duration: 10m

        # Experiment (A/B testing)
        - experiment:
            duration: 10m
            templates:
              - name: experiment-baseline
                specRef: stable
                weight: 50
              - name: experiment-canary
                specRef: canary
                weight: 50
            analyses:
              - name: experiment-analysis
                templateName: experiment-comparison

        # Final steps
        - setWeight: 100

Dynamic Pauses

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: dynamic-pause
spec:
  replicas: 5
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      steps:
        - setWeight: 20

        # Automatic pause (analysis-driven)
        - pause:
            duration: 5m

        # Progress to next step automatically if analysis passes
        - analysis:
            templates:
              - templateName: quick-check

        - setWeight: 50

        # Manual gate - requires explicit promotion
        - pause: {}

        # Note: Manual promotion command
        # kubectl argo rollouts promote myapp

Anti-Flapping Configuration

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: anti-flapping
spec:
  replicas: 10

  # Minimum time before health check
  minReadySeconds: 30

  # Maximum time to wait for rollout to progress
  progressDeadlineSeconds: 600

  # Abort on progress deadline
  progressDeadlineAbort: true

  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      maxSurge: 1
      maxUnavailable: 0

      steps:
        - setWeight: 25
        - pause:
            duration: 2m
        - setWeight: 50
        - pause:
            duration: 2m
        - setWeight: 75
        - pause:
            duration: 2m

Background Analysis Throughout Rollout

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: background-analysis
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      steps:
        - setWeight: 20
        - pause: {duration: 5m}
        - setWeight: 40
        - pause: {duration: 5m}
        - setWeight: 60
        - pause: {duration: 5m}
        - setWeight: 80
        - pause: {duration: 5m}

      # Background analysis runs continuously
      analysis:
        templates:
          - templateName: continuous-monitoring

        # Start analysis at step index 1
        startingStep: 1

        args:
          - name: service-name
            value: myapp
          - name: canary-hash
            valueFrom:
              podTemplateHashValue: Latest

Multi-Stage Analysis

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: multi-stage-analysis
spec:
  replicas: 20
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      maxSurge: "25%"
      maxUnavailable: 0

      steps:
        # Stage 1: Initial rollout with basic health checks
        - setWeight: 5
        - pause: {duration: 2m}
        - analysis:
            templates:
              - templateName: basic-health

        # Stage 2: Expand with error rate monitoring
        - setWeight: 20
        - pause: {duration: 5m}
        - analysis:
            templates:
              - templateName: error-rate-check
              - templateName: latency-check

        # Stage 3: Half traffic with comprehensive analysis
        - setWeight: 50
        - pause: {duration: 10m}
        - analysis:
            templates:
              - templateName: error-rate-check
              - templateName: latency-check
              - templateName: business-metrics

        # Stage 4: Manual approval gate
        - pause: {}

        # Stage 5: Final rollout
        - setWeight: 80
        - pause: {duration: 10m}
        - analysis:
            templates:
              - templateName: final-verification

      # Background analysis throughout
      analysis:
        templates:
          - templateName: continuous-alerts
        startingStep: 1

Automated Rollback Scenarios

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: auto-rollback
spec:
  replicas: 10

  # Abort if rollout doesn't progress
  progressDeadlineSeconds: 600
  progressDeadlineAbort: true

  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      maxSurge: "25%"
      maxUnavailable: 0

      # Automatically abort on failed analysis
      abortScaleDownDelaySeconds: 30

      steps:
        - setWeight: 10
        - pause: {duration: 2m}

        # Analysis will auto-rollback if metrics fail
        - analysis:
            templates:
              - templateName: critical-metrics
            args:
              - name: service
                value: myapp
              - name: canary-hash
                valueFrom:
                  podTemplateHashValue: Latest

        - setWeight: 30
        - pause: {duration: 5m}

        # Multiple analysis templates - ANY failure triggers rollback
        - analysis:
            templates:
              - templateName: error-rate
              - templateName: latency
              - templateName: saturation
            args:
              - name: service
                value: myapp

        - setWeight: 50
        - pause: {}

---
# Analysis template with strict failure conditions
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: critical-metrics
spec:
  args:
    - name: service
    - name: canary-hash

  metrics:
    - name: error-rate
      interval: 1m
      count: 5
      # Allow only 1 consecutive failure
      failureLimit: 1
      # Inconclusive results treated as failures
      inconclusiveLimit: 1

      successCondition: result < 0.01
      # Hard failure threshold
      failureCondition: result >= 0.05

      provider:
        prometheus:
          address: http://prometheus:9090
          query: |
            sum(rate(http_requests_total{
              service="{{args.service}}",
              rollouts_pod_template_hash="{{args.canary-hash}}",
              status=~"5.."
            }[5m]))
            /
            sum(rate(http_requests_total{
              service="{{args.service}}",
              rollouts_pod_template_hash="{{args.canary-hash}}"
            }[5m]))

Integration with Ingress Controllers

Integrating Argo Rollouts with ingress controllers for traffic splitting.

Key Concepts

  • Traffic Routing: Ingress controller routes traffic based on weights
  • Stable/Canary Services: Separate services for each version
  • Header-Based Routing: Route based on headers for testing
  • Annotation Management: Rollout controller updates ingress annotations
Traffic Management80%20%InternetIngress ControllerTraffic SplitStable ServiceCanary ServiceStable Pods v1Canary Pods v2Rollout ControllerUpdate WeightsTraffic Management80%20%InternetIngress ControllerTraffic SplitStable ServiceCanary ServiceStable Pods v1Canary Pods v2Rollout ControllerUpdate Weights

NGINX Ingress

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-nginx
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0
          ports:
            - containerPort: 8080

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        nginx:
          # Name of main ingress
          stableIngress: myapp-ingress

          # Optional: Annotation prefix (default: nginx.ingress.kubernetes.io)
          annotationPrefix: nginx.ingress.kubernetes.io

          # Optional: Additional ingress annotations
          additionalIngressAnnotations:
            canary-by-header: X-Canary
            canary-by-header-value: "true"

      steps:
        - setWeight: 10
        - pause: {duration: 5m}
        - setWeight: 25
        - pause: {duration: 5m}
        - setWeight: 50
        - pause: {duration: 10m}
        - setWeight: 75
        - pause: {duration: 10m}

---
# Stable service
apiVersion: v1
kind: Service
metadata:
  name: myapp-stable
spec:
  selector:
    app: myapp
  ports:
    - port: 80
      targetPort: 8080

---
# Canary service
apiVersion: v1
kind: Service
metadata:
  name: myapp-canary
spec:
  selector:
    app: myapp
  ports:
    - port: 80
      targetPort: 8080

---
# Ingress
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: myapp-ingress
  annotations:
    nginx.ingress.kubernetes.io/rewrite-target: /
spec:
  ingressClassName: nginx
  rules:
    - host: myapp.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: myapp-stable
                port:
                  number: 80

ALB (AWS Application Load Balancer)

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-alb
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0
          ports:
            - containerPort: 8080

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        alb:
          # Name of the ingress
          ingress: myapp-ingress

          # Service port
          servicePort: 80

          # Optional: Root service (for multiple paths)
          rootService: myapp-root

          # Optional: Sticky sessions
          stickinessConfig:
            enabled: true
            durationSeconds: 3600

      steps:
        - setWeight: 10
        - pause: {duration: 5m}
        - setWeight: 30
        - pause: {duration: 10m}
        - setWeight: 50
        - pause: {}

---
# ALB Ingress
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: myapp-ingress
  annotations:
    alb.ingress.kubernetes.io/scheme: internet-facing
    alb.ingress.kubernetes.io/target-type: ip
    alb.ingress.kubernetes.io/healthcheck-path: /health
spec:
  ingressClassName: alb
  rules:
    - host: myapp.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: myapp-stable
                port:
                  number: 80

Traefik

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-traefik
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0
          ports:
            - containerPort: 8080

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        traefik:
          # Traefik weighted round robin
          weightedTraefikServiceName: myapp-weighted

      steps:
        - setWeight: 20
        - pause: {duration: 5m}
        - setWeight: 50
        - pause: {duration: 10m}
        - setWeight: 80
        - pause: {duration: 10m}

---
# Traefik IngressRoute
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
  name: myapp-route
spec:
  entryPoints:
    - web
  routes:
    - match: Host(`myapp.example.com`)
      kind: Rule
      services:
        - name: myapp-weighted
          port: 80

---
# Traefik weighted service (managed by Rollout)
apiVersion: traefik.io/v1alpha1
kind: TraefikService
metadata:
  name: myapp-weighted
spec:
  weighted:
    services:
      - name: myapp-stable
        port: 80
        weight: 100
      - name: myapp-canary
        port: 80
        weight: 0

Header-Based Routing

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: header-routing
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        nginx:
          stableIngress: myapp-ingress

          # Enable header-based routing
          additionalIngressAnnotations:
            # Users with X-Canary: always header get canary
            canary-by-header: X-Canary
            canary-by-header-value: "always"

      steps:
        # Start with header-only routing (0% weight)
        - setWeight: 0
        - pause: {duration: 10m}  # Test with header

        # Begin percentage-based rollout
        - setWeight: 10
        - pause: {duration: 5m}
        - setWeight: 30
        - pause: {duration: 5m}
        - setWeight: 50
        - pause: {}

# Test with: curl -H "X-Canary: always" https://myapp.example.com

Integration with Service Meshes

Integrating Argo Rollouts with service meshes for advanced traffic management.

Key Concepts

  • Virtual Service: Service mesh routing rules
  • Destination Rule: Service mesh subset definitions
  • Traffic Splitting: Mesh-level traffic control
  • Mutual TLS: Secure service-to-service communication
Service Mesh Integration80%20%ClientVirtual ServiceTraffic SplitStable SubsetCanary SubsetStable Pods v1Canary Pods v2Rollout ControllerUpdate WeightsService Mesh Integration80%20%ClientVirtual ServiceTraffic SplitStable SubsetCanary SubsetStable Pods v1Canary Pods v2Rollout ControllerUpdate Weights

Istio Integration

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-istio
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
        version: v2  # Version label for Istio
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0
          ports:
            - containerPort: 8080

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        istio:
          # Virtual service to manage
          virtualService:
            name: myapp-vsvc

            # Optional: Specific routes to update
            routes:
              - primary

          # Optional: Destination rule with subsets
          destinationRule:
            name: myapp-destrule
            canarySubsetName: canary
            stableSubsetName: stable

      steps:
        - setWeight: 10
        - pause: {duration: 5m}
        - setWeight: 25
        - pause: {duration: 5m}
        - setWeight: 50
        - pause: {duration: 10m}
        - setWeight: 75
        - pause: {duration: 10m}

---
# Istio Virtual Service
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: myapp-vsvc
spec:
  hosts:
    - myapp.example.com
  gateways:
    - myapp-gateway
  http:
    - name: primary
      route:
        - destination:
            host: myapp-stable
            subset: stable
          weight: 100
        - destination:
            host: myapp-canary
            subset: canary
          weight: 0

---
# Istio Destination Rule
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
  name: myapp-destrule
spec:
  host: myapp-stable
  subsets:
    - name: stable
      labels:
        version: v1
    - name: canary
      labels:
        version: v2
  trafficPolicy:
    tls:
      mode: ISTIO_MUTUAL

---
# Gateway
apiVersion: networking.istio.io/v1
kind: Gateway
metadata:
  name: myapp-gateway
spec:
  selector:
    istio: ingressgateway
  servers:
    - port:
        number: 80
        name: http
        protocol: HTTP
      hosts:
        - myapp.example.com

Linkerd Integration (SMI)

Note: Linkerd 2.11+ no longer ships the SMI TrafficSplit CRD. On modern Linkerd you must install the SMI extension/CRD separately (or use a different mesh) for this smi traffic routing to work.

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: myapp-linkerd
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
      annotations:
        # Linkerd proxy injection
        linkerd.io/inject: enabled
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0
          ports:
            - containerPort: 8080

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        smi:
          # SMI TrafficSplit resource
          trafficSplitName: myapp-trafficsplit

          # Optional: Root service
          rootService: myapp-root

      steps:
        - setWeight: 10
        - pause: {duration: 5m}
        - setWeight: 30
        - pause: {duration: 10m}
        - setWeight: 50
        - pause: {}

---
# SMI Traffic Split (managed by Rollout)
apiVersion: split.smi-spec.io/v1alpha1
kind: TrafficSplit
metadata:
  name: myapp-trafficsplit
spec:
  service: myapp-root
  backends:
    - service: myapp-stable
      weight: 100
    - service: myapp-canary
      weight: 0

---
# Root service
apiVersion: v1
kind: Service
metadata:
  name: myapp-root
spec:
  selector:
    app: myapp
  ports:
    - port: 80
      targetPort: 8080

Istio with Header-Based Routing

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: advanced-istio
spec:
  replicas: 10
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
        version: v2
    spec:
      containers:
        - name: myapp
          image: myapp:v2.0.0

  strategy:
    canary:
      stableService: myapp-stable
      canaryService: myapp-canary

      trafficRouting:
        istio:
          virtualService:
            name: myapp-vsvc
            routes:
              - primary

      steps:
        # Start with header-based routing
        - setHeaderRoute:
            name: test-users
            match:
              - headerName: X-Test-User
                headerValue:
                  exact: "true"
        - pause: {duration: 10m}

        # Begin percentage rollout
        - setWeight: 10
        - pause: {duration: 5m}
        - setWeight: 25
        - pause: {duration: 5m}
        - setWeight: 50
        - pause: {}

---
# Istio Virtual Service with header matching
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: myapp-vsvc
spec:
  hosts:
    - myapp.example.com
  http:
    # Header-based route (managed by Rollout)
    - name: test-users
      match:
        - headers:
            X-Test-User:
              exact: "true"
      route:
        - destination:
            host: myapp-canary
          weight: 100

    # Primary route (managed by Rollout)
    - name: primary
      route:
        - destination:
            host: myapp-stable
          weight: 100
        - destination:
            host: myapp-canary
          weight: 0

Quick Reference

Common Commands

Command Description
kubectl argo rollouts list List all rollouts
kubectl argo rollouts get rollout <name> Get rollout status
kubectl argo rollouts get rollout <name> --watch Watch rollout progress
kubectl argo rollouts promote <name> Promote to next step
kubectl argo rollouts promote <name> --full Skip to end
kubectl argo rollouts pause <name> Pause rollout
kubectl argo rollouts resume <name> Resume paused rollout
kubectl argo rollouts abort <name> Abort rollout
kubectl argo rollouts retry rollout <name> Retry failed rollout
kubectl argo rollouts undo <name> Rollback to previous
kubectl argo rollouts undo <name> --to-revision=N Rollback to revision N
kubectl argo rollouts set image <name> <container>=<image> Update image
kubectl argo rollouts restart <name> Restart rollout
kubectl argo rollouts status <name> Get rollout status
kubectl argo rollouts dashboard Launch dashboard

Rollout Status Values

Status Description
Healthy All replicas healthy and at desired version
Progressing Rollout in progress
Degraded Rollout has failed or exceeded deadline
Paused Rollout paused waiting for promotion

Analysis Result Values

Result Description
Successful Analysis passed all metrics
Failed Analysis failed metrics threshold
Error Analysis encountered an error
Running Analysis currently executing
Pending Analysis not yet started
Inconclusive Cannot determine result

Metrics Providers

Provider Use Case
Prometheus Open-source metrics and alerting
Datadog Cloud monitoring and analytics
New Relic Application performance monitoring
Wavefront Cloud-native monitoring
Kayenta Statistical analysis engine
Web HTTP endpoint checks
Job Kubernetes job-based tests
CloudWatch AWS metrics
Graphite Time-series database

Traffic Routing Options

Method Description
Nginx Ingress NGINX-based traffic splitting
ALB AWS Application Load Balancer
Istio Service mesh with virtual services
Linkerd Lightweight service mesh (SMI)
Traefik Cloud-native edge router
App Mesh AWS service mesh
Ambassador API Gateway with traffic control

Step Types Reference

Step Type Description Example
setWeight Set traffic weight percentage setWeight: 20
pause Pause with duration or indefinitely pause: {duration: 5m} or pause: {}
setCanaryScale Set canary replica count setCanaryScale: {replicas: 3}
analysis Run inline analysis analysis: {templates: [...]}
experiment Run A/B test experiment experiment: {duration: 10m}
setHeaderRoute Route by header (Istio) setHeaderRoute: {name: test}

Common Issues and Solutions

Issue Cause Solution
Rollout stuck in Progressing Analysis timeout or failure Check AnalysisRun: kubectl get analysisrun -n <namespace>
Traffic not splitting Ingress/mesh misconfiguration Verify service names and traffic routing config match
Analysis always fails Incorrect query or threshold Test query directly against metrics provider
Pods not scaling Resource constraints Check cluster capacity: kubectl describe nodes
Rollout not starting Deployment hasn't changed Update pod template to trigger new revision
Cannot promote No manual pause step Ensure strategy includes pause: {} step
Service mesh integration fails Missing VirtualService/DestinationRule Create required mesh resources first
Analysis provider timeout Network or authentication issues Check provider connectivity and credentials
Rollback not working No previous revision Ensure revisionHistoryLimit > 1
Dashboard not showing rollout Namespace filtering Check dashboard is monitoring correct namespace
Canary service has no endpoints Selector doesn't match pods Verify service selector matches pod labels
Analysis inconclusive No data from provider Check metric exists and query returns data

Debugging Commands

# View rollout events
kubectl describe rollout <name> -n <namespace>

# Check analysis runs
kubectl get analysisrun -n <namespace>
kubectl describe analysisrun <name> -n <namespace>

# View controller logs
kubectl logs -n argo-rollouts deployment/argo-rollouts -f

# Check ReplicaSet status
kubectl get rs -l app=<app-name> -n <namespace>

# View pod status and labels
kubectl get pods -l app=<app-name> -n <namespace> --show-labels

# Check traffic routing config
kubectl get ingress <name> -n <namespace> -o yaml
kubectl get virtualservice <name> -n <namespace> -o yaml

# Export rollout for inspection
kubectl get rollout <name> -n <namespace> -o yaml > rollout-debug.yaml

# Validate analysis template
kubectl apply --dry-run=client -f analysis-template.yaml

# Check service endpoints
kubectl get endpoints <service-name> -n <namespace>

# View rollout in JSON format
kubectl argo rollouts get rollout <name> -n <namespace> -o json

Troubleshooting Analysis

# List all analysis runs
kubectl get analysisrun -n <namespace>

# Get analysis run details
kubectl describe analysisrun <name> -n <namespace>

# View analysis run status
kubectl get analysisrun <name> -n <namespace> -o jsonpath='{.status}'

# Check metric results
kubectl get analysisrun <name> -n <namespace> -o jsonpath='{.status.metricResults}'

# Test Prometheus query manually
kubectl run -it --rm debug --image=curlimages/curl --restart=Never -- \
  curl "http://prometheus:9090/api/v1/query?query=<encoded-query>"

# Delete failed analysis run to retry
kubectl delete analysisrun <name> -n <namespace>

# Watch analysis run progress
kubectl get analysisrun <name> -n <namespace> -w

# Check analysis template args
kubectl get analysisrun <name> -n <namespace> -o jsonpath='{.spec.args}'

Troubleshooting Traffic Routing

# NGINX Ingress: Check canary ingress creation
kubectl get ingress -n <namespace> | grep canary

# NGINX Ingress: View canary annotations
kubectl get ingress <name>-<hash>-canary -n <namespace> -o yaml

# Istio: Check VirtualService weights
kubectl get virtualservice <name> -n <namespace> -o yaml | grep -A 5 weight

# Linkerd: Check TrafficSplit status
kubectl get trafficsplit <name> -n <namespace> -o yaml

# Check service selectors
kubectl get svc <name> -n <namespace> -o yaml | grep -A 5 selector

# Verify pod labels match service
kubectl get pods -l app=<name> -n <namespace> --show-labels

Related Topics

The following topics complement Argo Rollouts and are commonly used together in progressive delivery workflows:

  1. ArgoCD - GitOps continuous delivery that integrates seamlessly with Argo Rollouts for deployment automation
  2. Istio - Service mesh providing advanced traffic management capabilities for canary deployments
  3. Prometheus - Metrics collection and querying for automated analysis during rollouts
  4. Kubernetes - Understanding Kubernetes services, deployments, and networking is essential for Rollouts
  5. Grafana - Visualising rollout metrics and creating dashboards for deployment monitoring
  6. NGINX Ingress - Popular ingress controller for traffic splitting during canary deployments