Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Envoy Proxy

Comprehensive guide to Envoy, a high-performance L7 proxy and communication bus for microservices.

Envoy Proxy

Comprehensive guide to Envoy, a high-performance L7 proxy and communication bus for microservices.

Overview

Envoy is a modern, high-performance edge and service proxy designed for cloud-native applications. Originally built at Lyft, Envoy provides advanced load balancing, observability, and security features for microservices architectures. It operates as both a transparent proxy and a programmable edge proxy, forming the data plane for service meshes like Istio and acting as an API gateway or standalone load balancer.

Key characteristics:

  • Out-of-process architecture (sidecar pattern)
  • L3/L4 and L7 filter architecture
  • HTTP/2 and gRPC native support
  • Advanced load balancing (active/passive health checks, circuit breaking, rate limiting)
  • Rich observability (stats, logging, tracing)
  • Dynamic configuration via xDS APIs
Upstream ServicesEnvoy ProxyListener :80Filter ChainHTTP ConnectionManagerHTTP FiltersRouter FilterCluster: backendService Instance 1Service Instance 2Service Instance 3ClientUpstream ServicesEnvoy ProxyListener :80Filter ChainHTTP ConnectionManagerHTTP FiltersRouter FilterCluster: backendService Instance 1Service Instance 2Service Instance 3Client

Request Flow

Understanding how Envoy processes requests is crucial for effective configuration.

EndpointClusterRouterHTTPFiltersFilterChainListenerClientEndpointClusterRouterHTTPFiltersFilterChainListenerClientIncoming RequestMatch Filter ChainApply HTTP Filtersext_authz, CORS, etc.Route MatchingSelect ClusterLoad BalanceResponseResponseResponse FiltersFinal ResponseEndpointClusterRouterHTTPFiltersFilterChainListenerClientEndpointClusterRouterHTTPFiltersFilterChainListenerClientIncoming RequestMatch Filter ChainApply HTTP Filtersext_authz, CORS, etc.Route MatchingSelect ClusterLoad BalanceResponseResponseResponse FiltersFinal Response

Listeners and Routes

Listeners define how Envoy accepts incoming connections, whilst routes determine where traffic is forwarded.

Key Concepts

Listener: Binds to a network address and port, accepting downstream connections. Each listener has a filter chain that processes the connection.

Filter Chain: Ordered list of network filters (L3/L4) and HTTP filters (L7) applied to connections.

Route Configuration: Defines how to match requests and forward them to clusters based on domains, paths, headers, and other criteria.

Virtual Hosts: Group routes by domain name, allowing multiple services on a single listener.

Basic Listener Configuration

static_resources:
  listeners:
  - name: main_listener
    address:
      socket_address:
        address: 0.0.0.0
        port_value: 8080
    filter_chains:
    - filters:
      - name: envoy.filters.network.http_connection_manager
        typed_config:
          "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
          stat_prefix: ingress_http
          codec_type: AUTO
          route_config:
            name: local_route
            virtual_hosts:
            - name: backend
              domains: ["*"]
              routes:
              - match:
                  prefix: "/"
                route:
                  cluster: backend_cluster
          http_filters:
          - name: envoy.filters.http.router
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

Advanced Routing

# Path-based routing with regex
route_config:
  name: local_route
  virtual_hosts:
  - name: api_backend
    domains: ["api.example.com", "api.example.com:*"]
    routes:
    # Exact match
    - match:
        path: "/health"
      route:
        cluster: health_cluster

    # Prefix match
    - match:
        prefix: "/api/v1/"
      route:
        cluster: api_v1_cluster
        prefix_rewrite: "/"

    # Regex match
    - match:
        safe_regex:
          regex: "^/api/v[0-9]+/users/.*"
      route:
        cluster: user_service

    # Header-based routing
    - match:
        prefix: "/api/"
        headers:
        - name: "x-version"
          string_match:
            exact: "v2"
      route:
        cluster: api_v2_cluster

Traffic Splitting

# Weighted routing for canary deployments
routes:
- match:
    prefix: "/api"
  route:
    weighted_clusters:
      clusters:
      - name: api_stable
        weight: 90
      - name: api_canary
        weight: 10
    total_weight: 100

Virtual Host Configuration

virtual_hosts:
- name: service_a
  domains: ["service-a.internal"]
  routes:
  - match:
      prefix: "/"
    route:
      cluster: service_a_cluster
      timeout: 5s
      retry_policy:
        retry_on: "5xx"
        num_retries: 3

- name: service_b
  domains: ["service-b.internal"]
  routes:
  - match:
      prefix: "/"
    route:
      cluster: service_b_cluster

Clusters

Clusters define groups of upstream endpoints that Envoy can route traffic to.

Key Concepts

Cluster: Logical group of upstream hosts that Envoy connects to. Clusters define load balancing policy, health checking, and connection settings.

Endpoint: Individual upstream host (IP:port) within a cluster.

Load Balancing: Algorithm for selecting endpoints (ROUND_ROBIN, LEAST_REQUEST, RANDOM, RING_HASH, MAGLEV).

Circuit Breaking: Limits concurrent connections and requests to prevent cascading failures.

Basic Cluster Configuration

static_resources:
  clusters:
  - name: backend_cluster
    type: STRICT_DNS
    lb_policy: ROUND_ROBIN
    load_assignment:
      cluster_name: backend_cluster
      endpoints:
      - lb_endpoints:
        - endpoint:
            address:
              socket_address:
                address: backend.example.com
                port_value: 8080

Cluster Discovery Types

# STATIC: Explicitly defined endpoints
- name: static_cluster
  type: STATIC
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: static_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: 10.0.1.10
              port_value: 8080
      - endpoint:
          address:
            socket_address:
              address: 10.0.1.11
              port_value: 8080

# STRICT_DNS: Resolve DNS name, all IPs become endpoints
- name: dns_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  dns_lookup_family: V4_ONLY
  load_assignment:
    cluster_name: dns_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: api.example.com
              port_value: 443

# LOGICAL_DNS: Resolve DNS on each new connection
- name: logical_dns_cluster
  type: LOGICAL_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: logical_dns_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: dynamic.example.com
              port_value: 8080

# EDS: Endpoint Discovery Service (xDS)
- name: eds_cluster
  type: EDS
  eds_cluster_config:
    eds_config:
      api_config_source:
        api_type: GRPC
        grpc_services:
        - envoy_grpc:
            cluster_name: xds_cluster

Load Balancing Policies

# Round robin
- name: rr_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: rr_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: backend.local
              port_value: 8080

# Least request (best for variable request duration)
- name: lr_cluster
  type: STRICT_DNS
  lb_policy: LEAST_REQUEST
  load_assignment:
    cluster_name: lr_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: backend.local
              port_value: 8080

# Ring hash (consistent hashing for sticky sessions)
- name: sticky_cluster
  type: STRICT_DNS
  lb_policy: RING_HASH
  ring_hash_lb_config:
    minimum_ring_size: 1024
  load_assignment:
    cluster_name: sticky_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: backend.local
              port_value: 8080

Connection Pool Settings

clusters:
- name: backend_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  # HTTP/1.1 connection pool
  circuit_breakers:
    thresholds:
    - priority: DEFAULT
      max_connections: 1024
      max_pending_requests: 1024
      max_requests: 1024
      max_retries: 3

  # HTTP/2 connection pool
  http2_protocol_options:
    max_concurrent_streams: 100

  # TCP connection settings
  upstream_connection_options:
    tcp_keepalive:
      keepalive_time: 300

Service Discovery and Health Checks

Envoy integrates with service discovery systems and actively monitors endpoint health.

Key Concepts

Active Health Checking: Envoy periodically sends health check requests to endpoints to determine availability.

Passive Health Checking: Envoy monitors actual traffic and marks endpoints unhealthy based on consecutive failures (outlier detection).

Endpoint Discovery Service (EDS): Dynamic API for discovering cluster endpoints, part of xDS protocol.

Active Health Checks

clusters:
- name: backend_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: backend_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: backend.local
              port_value: 8080

  # HTTP health check
  health_checks:
  - timeout: 1s
    interval: 10s
    unhealthy_threshold: 3
    healthy_threshold: 2
    http_health_check:
      path: "/health"
      expected_statuses:
      - start: 200
        end: 299
      request_headers_to_add:
      - header:
          key: "User-Agent"
          value: "envoy-health-checker"

TCP Health Checks

health_checks:
- timeout: 1s
  interval: 5s
  unhealthy_threshold: 2
  healthy_threshold: 2
  tcp_health_check:
    # send/receive payloads are hex-encoded bytes ("ping"/"pong")
    send:
      text: "70696E67"
    receive:
    - text: "706F6E67"

gRPC Health Checks

health_checks:
- timeout: 1s
  interval: 10s
  unhealthy_threshold: 3
  healthy_threshold: 2
  grpc_health_check:
    service_name: "my.service.v1.HealthService"

Outlier Detection (Passive Health Checking)

clusters:
- name: backend_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  outlier_detection:
    # Consecutive 5xx errors before ejection
    consecutive_5xx: 5
    # Consecutive gateway failures before ejection
    consecutive_gateway_failure: 3
    # Time interval for analysis
    interval: 10s
    # Base ejection time (increases exponentially)
    base_ejection_time: 30s
    # Maximum ejection time
    max_ejection_time: 300s
    # Maximum percentage of hosts that can be ejected
    max_ejection_percent: 50
    # Success rate to enforce (95%)
    success_rate_minimum_hosts: 5
    success_rate_request_volume: 100
    success_rate_stdev_factor: 1900

Dynamic Endpoint Discovery

# Configure EDS cluster
clusters:
- name: backend_cluster
  type: EDS
  lb_policy: ROUND_ROBIN
  eds_cluster_config:
    eds_config:
      resource_api_version: V3
      api_config_source:
        api_type: GRPC
        transport_api_version: V3
        grpc_services:
        - envoy_grpc:
            cluster_name: xds_cluster

# XDS management server cluster
- name: xds_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  http2_protocol_options: {}
  load_assignment:
    cluster_name: xds_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: control-plane.local
              port_value: 18000

HTTP Filters

HTTP filters process L7 requests and responses in the filter chain.

Key Concepts

Filter Chain: Ordered list of filters applied to requests (and responses in reverse order).

Router Filter: Required terminal filter that forwards requests to upstream clusters.

Decoder Filters: Process requests from downstream (client → Envoy).

Encoder Filters: Process responses from upstream (Envoy → client).

Decoder/Encoder Filters: Process both requests and responses.

Common HTTP Filters

http_filters:
# CORS filter
- name: envoy.filters.http.cors
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.cors.v3.Cors

# Request size limit
- name: envoy.filters.http.buffer
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.buffer.v3.Buffer
    max_request_bytes: 5242880  # 5MB

# gRPC web support
- name: envoy.filters.http.grpc_web
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.grpc_web.v3.GrpcWeb

# Health check filter (responds without forwarding)
- name: envoy.filters.http.health_check
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.health_check.v3.HealthCheck
    pass_through_mode: false
    headers:
    - name: ":path"
      string_match:
        exact: "/healthz"

# Rate limiting
- name: envoy.filters.http.ratelimit
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.ratelimit.v3.RateLimit
    domain: backend_ratelimit
    failure_mode_deny: false
    rate_limit_service:
      grpc_service:
        envoy_grpc:
          cluster_name: ratelimit_cluster

# Router (must be last)
- name: envoy.filters.http.router
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

Header Manipulation

routes:
- match:
    prefix: "/api"
  route:
    cluster: backend_cluster
  # Header mutations sit at the route entry level (sibling of `route`)
  # Add/override request headers
  request_headers_to_add:
  - header:
      key: "x-custom-header"
      value: "my-value"
    append_action: OVERWRITE_IF_EXISTS_OR_ADD
  - header:
      key: "x-forwarded-proto"
      value: "https"

  # Remove request headers
  request_headers_to_remove:
  - "x-internal-header"

  # Add response headers
  response_headers_to_add:
  - header:
      key: "x-served-by"
      value: "envoy"

  # Remove response headers
  response_headers_to_remove:
  - "x-powered-by"

Fault Injection

routes:
- match:
    prefix: "/api"
  route:
    cluster: backend_cluster
  typed_per_filter_config:
    envoy.filters.http.fault:
      "@type": type.googleapis.com/envoy.extensions.filters.http.fault.v3.HTTPFault
      # Delay 50% of requests by 5s
      delay:
        percentage:
          numerator: 50
          denominator: HUNDRED
        fixed_delay: 5s
      # Abort 10% of requests with 503
      abort:
        percentage:
          numerator: 10
          denominator: HUNDRED
        http_status: 503

External Authorisation (ext_authz)

The ext_authz filter delegates authorisation decisions to an external service.

Key Concepts

External Authorisation: Offload authentication/authorisation to a separate service (e.g., OPA, custom auth service).

Check Request: Envoy sends request metadata to auth service before forwarding to upstream.

Check Response: Auth service returns OK (allow) or Denied (reject with status code).

Context Extensions: Pass additional metadata to the auth service.

HTTP ext_authz Configuration

http_filters:
- name: envoy.filters.http.ext_authz
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthz
    # Use HTTP auth service
    http_service:
      server_uri:
        uri: "http://auth-service:9191"
        cluster: auth_cluster
        timeout: 0.5s

      # Headers to send to auth service
      authorization_request:
        allowed_headers:
          patterns:
          - exact: "authorization"
          - exact: "cookie"
          - prefix: "x-auth-"

      # Headers to forward from auth response
      authorization_response:
        allowed_upstream_headers:
          patterns:
          - exact: "x-user-id"
          - exact: "x-user-roles"
        allowed_client_headers:
          patterns:
          - exact: "set-cookie"

      # Path prefix for auth check
      path_prefix: "/auth/check"

    # Behaviour on failure
    failure_mode_allow: false

    # Include request body in check (use carefully)
    with_request_body:
      max_request_bytes: 8192
      allow_partial_message: true

- name: envoy.filters.http.router
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

clusters:
- name: auth_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: auth_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: auth-service.local
              port_value: 9191

gRPC ext_authz Configuration

http_filters:
- name: envoy.filters.http.ext_authz
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthz
    # Use gRPC auth service
    grpc_service:
      envoy_grpc:
        cluster_name: auth_grpc_cluster
      timeout: 0.5s

    # Include peer certificate in check
    include_peer_certificate: true

    failure_mode_allow: false

clusters:
- name: auth_grpc_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  http2_protocol_options: {}
  load_assignment:
    cluster_name: auth_grpc_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: auth-grpc.local
              port_value: 9192

Per-Route ext_authz Override

routes:
- match:
    prefix: "/public"
  route:
    cluster: backend_cluster
  typed_per_filter_config:
    envoy.filters.http.ext_authz:
      "@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthzPerRoute
      # Disable auth for this route
      disabled: true

- match:
    prefix: "/admin"
  route:
    cluster: admin_cluster
  typed_per_filter_config:
    envoy.filters.http.ext_authz:
      "@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthzPerRoute
      check_settings:
        context_extensions:
          required_role: "admin"

TCP Filters

TCP filters operate at L3/L4 for non-HTTP protocols.

Key Concepts

Network Filter: Operates on raw TCP streams before L7 processing.

TCP Proxy: Terminal filter that forwards TCP connections to upstream clusters.

SNI-based Routing: Route based on TLS Server Name Indication without decrypting.

TCP Proxy Configuration

listeners:
- name: tcp_listener
  address:
    socket_address:
      address: 0.0.0.0
      port_value: 3306
  filter_chains:
  - filters:
    - name: envoy.filters.network.tcp_proxy
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy
        stat_prefix: mysql_tcp
        cluster: mysql_cluster
        # Idle timeout
        idle_timeout: 3600s
        # Access logging
        access_log:
        - name: envoy.access_loggers.file
          typed_config:
            "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
            path: "/var/log/envoy/tcp-access.log"

clusters:
- name: mysql_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: mysql_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: mysql.local
              port_value: 3306

SNI-based Routing

listeners:
- name: tls_listener
  address:
    socket_address:
      address: 0.0.0.0
      port_value: 443
  filter_chains:
  # Route for service-a.example.com
  - filter_chain_match:
      server_names: ["service-a.example.com"]
    filters:
    - name: envoy.filters.network.tcp_proxy
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy
        stat_prefix: service_a
        cluster: service_a_cluster

  # Route for service-b.example.com
  - filter_chain_match:
      server_names: ["service-b.example.com"]
    filters:
    - name: envoy.filters.network.tcp_proxy
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy
        stat_prefix: service_b
        cluster: service_b_cluster

  # Default route
  - filters:
    - name: envoy.filters.network.tcp_proxy
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy
        stat_prefix: default
        cluster: default_cluster

MongoDB Proxy Filter

filter_chains:
- filters:
  - name: envoy.filters.network.mongo_proxy
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.network.mongo_proxy.v3.MongoProxy
      stat_prefix: mongo
      access_log: "/var/log/envoy/mongo-access.log"

  - name: envoy.filters.network.tcp_proxy
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy
      stat_prefix: mongo_tcp
      cluster: mongo_cluster

Traffic Shaping

Control traffic flow with timeouts, retries, and circuit breakers.

Key Concepts

Timeouts: Limits on how long Envoy waits for responses at various stages.

Retries: Automatic retry logic for failed requests based on configurable conditions.

Circuit Breakers: Prevent cascading failures by limiting concurrent connections and requests.

Rate Limiting: Control request rate to protect upstream services.

Timeouts

routes:
- match:
    prefix: "/api"
  route:
    cluster: backend_cluster
    # Total request timeout (includes retries)
    timeout: 15s
    # Idle timeout for streaming
    idle_timeout: 60s
    # Per-try timeout (for each retry attempt)
    max_stream_duration:
      max_stream_duration: 5s
      grpc_timeout_header_max: 5s

clusters:
- name: backend_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  # Connection timeout
  connect_timeout: 2s
  # DNS refresh rate
  dns_refresh_rate: 30s
  load_assignment:
    cluster_name: backend_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: backend.local
              port_value: 8080

Retry Policy

routes:
- match:
    prefix: "/api"
  route:
    cluster: backend_cluster
    retry_policy:
      # Conditions to trigger retry
      retry_on: "5xx,reset,connect-failure,refused-stream"
      # Number of retry attempts
      num_retries: 3
      # Per-try timeout
      per_try_timeout: 2s
      # Per-try idle timeout
      per_try_idle_timeout: 1s
      # Retry budget (prevent retry storms)
      retry_host_predicate:
      - name: envoy.retry_host_predicates.previous_hosts
        typed_config:
          "@type": type.googleapis.com/envoy.extensions.retry.host.previous_hosts.v3.PreviousHostsPredicate
      # Exponential backoff
      retry_back_off:
        base_interval: 0.1s
        max_interval: 1s
      # Host selection retry
      host_selection_retry_max_attempts: 5

Advanced Retry Configuration

# Retry only on specific HTTP status codes
retry_policy:
  retry_on: "retriable-status-codes"
  retriable_status_codes:
  - 502
  - 503
  - 504
  num_retries: 3

# Retry based on headers
retry_policy:
  retry_on: "retriable-headers"
  retriable_headers:
  - name: "x-retry"
    string_match:
      exact: "true"
  num_retries: 2

# Rate-limited retry budget
retry_policy:
  retry_on: "5xx"
  num_retries: 3
  rate_limited_retry_back_off:
    reset_headers:
    - name: "retry-after"
    max_interval: 300s

Circuit Breakers

clusters:
- name: backend_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  circuit_breakers:
    thresholds:
    # Default priority
    - priority: DEFAULT
      # Maximum concurrent connections to all endpoints
      max_connections: 1024
      # Maximum pending requests waiting for connections
      max_pending_requests: 1024
      # Maximum concurrent requests across all connections
      max_requests: 1024
      # Maximum concurrent retries
      max_retries: 3
      # Track request volume for outlier detection
      track_remaining: true

    # High priority traffic (separate limits)
    - priority: HIGH
      max_connections: 2048
      max_pending_requests: 2048
      max_requests: 2048
      max_retries: 5

  load_assignment:
    cluster_name: backend_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: backend.local
              port_value: 8080

Local Rate Limiting

http_filters:
# Local rate limit (per Envoy instance)
- name: envoy.filters.http.local_ratelimit
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.local_ratelimit.v3.LocalRateLimit
    stat_prefix: http_local_rate_limiter
    token_bucket:
      max_tokens: 100
      tokens_per_fill: 10
      fill_interval: 1s
    filter_enabled:
      runtime_key: local_rate_limit_enabled
      default_value:
        numerator: 100
        denominator: HUNDRED
    filter_enforced:
      runtime_key: local_rate_limit_enforced
      default_value:
        numerator: 100
        denominator: HUNDRED
    response_headers_to_add:
    - append_action: OVERWRITE_IF_EXISTS_OR_ADD
      header:
        key: x-local-rate-limit
        value: 'true'

- name: envoy.filters.http.router
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

Observability

Envoy provides rich telemetry through access logs, metrics, and distributed tracing.

Key Concepts

Access Logs: Detailed logs of each request/connection with customisable format.

Stats: Metrics exposed via admin interface and StatsD/Prometheus.

Distributed Tracing: Integration with Zipkin, Jaeger, Datadog, etc.

Admin Interface: HTTP endpoint for configuration, stats, and runtime management.

Access Logging

# HTTP access logs
http_connection_manager:
  stat_prefix: ingress_http
  codec_type: AUTO
  route_config:
    name: local_route
    virtual_hosts:
    - name: backend
      domains: ["*"]
      routes:
      - match:
          prefix: "/"
        route:
          cluster: backend_cluster

  # File access log with JSON format
  access_log:
  - name: envoy.access_loggers.file
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
      path: "/var/log/envoy/access.log"
      json_format:
        timestamp: "%START_TIME%"
        method: "%REQ(:METHOD)%"
        path: "%REQ(X-ENVOY-ORIGINAL-PATH?:PATH)%"
        protocol: "%PROTOCOL%"
        response_code: "%RESPONSE_CODE%"
        response_flags: "%RESPONSE_FLAGS%"
        bytes_received: "%BYTES_RECEIVED%"
        bytes_sent: "%BYTES_SENT%"
        duration: "%DURATION%"
        upstream_service_time: "%RESP(X-ENVOY-UPSTREAM-SERVICE-TIME)%"
        forwarded_for: "%REQ(X-FORWARDED-FOR)%"
        user_agent: "%REQ(USER-AGENT)%"
        request_id: "%REQ(X-REQUEST-ID)%"
        authority: "%REQ(:AUTHORITY)%"
        upstream_host: "%UPSTREAM_HOST%"
        upstream_cluster: "%UPSTREAM_CLUSTER%"

  # Custom text format
  - name: envoy.access_loggers.file
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
      path: "/dev/stdout"
      format: "[%START_TIME%] \"%REQ(:METHOD)% %REQ(X-ENVOY-ORIGINAL-PATH?:PATH)% %PROTOCOL%\" %RESPONSE_CODE% %RESPONSE_FLAGS% %BYTES_RECEIVED% %BYTES_SENT% %DURATION% \"%REQ(X-FORWARDED-FOR)%\" \"%REQ(USER-AGENT)%\" \"%REQ(X-REQUEST-ID)%\" \"%REQ(:AUTHORITY)%\" \"%UPSTREAM_HOST%\"\n"

  http_filters:
  - name: envoy.filters.http.router
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

Conditional Access Logging

access_log:
# Only log errors (4xx/5xx)
- name: envoy.access_loggers.file
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
    path: "/var/log/envoy/error.log"
  filter:
    status_code_filter:
      comparison:
        op: GE
        value:
          default_value: 400
          runtime_key: access_log_filter_http_code

# Only log slow requests (>1s)
- name: envoy.access_loggers.file
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
    path: "/var/log/envoy/slow.log"
  filter:
    duration_filter:
      comparison:
        op: GE
        value:
          default_value: 1000
          runtime_key: access_log_filter_duration_ms

gRPC Access Log Service

access_log:
- name: envoy.access_loggers.http_grpc
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.access_loggers.grpc.v3.HttpGrpcAccessLogConfig
    common_config:
      log_name: "http_access_log"
      transport_api_version: V3
      grpc_service:
        envoy_grpc:
          cluster_name: access_log_cluster
    additional_request_headers_to_log:
    - "x-request-id"
    - "x-user-id"
    additional_response_headers_to_log:
    - "x-response-time"

clusters:
- name: access_log_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  http2_protocol_options: {}
  load_assignment:
    cluster_name: access_log_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: log-collector.local
              port_value: 9001

Statistics Configuration

stats_sinks:
# Prometheus metrics
- name: envoy.stat_sinks.metrics_service
  typed_config:
    "@type": type.googleapis.com/envoy.config.metrics.v3.MetricsServiceConfig
    transport_api_version: V3
    grpc_service:
      envoy_grpc:
        cluster_name: metrics_cluster

# StatsD integration
- name: envoy.stat_sinks.statsd
  typed_config:
    "@type": type.googleapis.com/envoy.config.metrics.v3.StatsdSink
    tcp_cluster_name: statsd_cluster
    prefix: "envoy"

# DogStatsD (Datadog)
- name: envoy.stat_sinks.dog_statsd
  typed_config:
    "@type": type.googleapis.com/envoy.config.metrics.v3.DogStatsdSink
    address:
      socket_address:
        address: 127.0.0.1
        port_value: 8125
    prefix: "envoy"

# Statistics configuration
stats_config:
  stats_tags:
  - tag_name: "cluster_name"
    regex: "^cluster\\.((.+?)\\.)"
  - tag_name: "listener_address"
    regex: "^listener\\.((.+?)\\.)"
  use_all_default_tags: true
  stats_matcher:
    inclusion_list:
      patterns:
      - prefix: "cluster.backend"
      - prefix: "http"

Distributed Tracing

# HTTP connection manager with tracing
http_connection_manager:
  stat_prefix: ingress_http
  codec_type: AUTO

  # Tracing configuration
  tracing:
    provider:
      name: envoy.tracers.zipkin
      typed_config:
        "@type": type.googleapis.com/envoy.config.trace.v3.ZipkinConfig
        collector_cluster: zipkin
        collector_endpoint: "/api/v2/spans"
        collector_endpoint_version: HTTP_JSON
        trace_id_128bit: true
        shared_span_context: false

    # Sampling configuration
    random_sampling:
      value: 100.0  # 100% sampling

    # Tag custom headers
    custom_tags:
    - tag: "user_id"
      request_header:
        name: "x-user-id"
    - tag: "environment"
      literal:
        value: "production"

  route_config:
    name: local_route
    virtual_hosts:
    - name: backend
      domains: ["*"]
      routes:
      - match:
          prefix: "/"
        route:
          cluster: backend_cluster
        decorator:
          operation: "backend_operation"

  http_filters:
  - name: envoy.filters.http.router
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router
      start_child_span: true

clusters:
- name: zipkin
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: zipkin
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: zipkin.local
              port_value: 9411

OpenTelemetry Tracing

The OpenCensus tracer was removed from Envoy; use the OpenTelemetry tracer to export to Jaeger, Tempo, or any OTLP collector (Jaeger ingests OTLP natively).

tracing:
  provider:
    name: envoy.tracers.opentelemetry
    typed_config:
      "@type": type.googleapis.com/envoy.config.trace.v3.OpenTelemetryConfig
      service_name: "envoy-proxy"
      # OTLP/gRPC exporter -> collector cluster (port 4317)
      grpc_service:
        envoy_grpc:
          cluster_name: opentelemetry_collector
        timeout: 0.25s

clusters:
- name: opentelemetry_collector
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  http2_protocol_options: {}
  load_assignment:
    cluster_name: opentelemetry_collector
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: otel-collector.local
              port_value: 4317

Admin Interface

admin:
  # Admin interface configuration
  address:
    socket_address:
      address: 0.0.0.0
      port_value: 9901

  # Access log for admin requests
  access_log:
  - name: envoy.access_loggers.file
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
      path: "/var/log/envoy/admin-access.log"

Common admin endpoints:

# View configuration
curl http://localhost:9901/config_dump

# View stats (Prometheus format)
curl http://localhost:9901/stats/prometheus

# View stats (JSON)
curl http://localhost:9901/stats?format=json

# View clusters status
curl http://localhost:9901/clusters

# View listeners
curl http://localhost:9901/listeners

# Health check
curl http://localhost:9901/ready
curl http://localhost:9901/server_info

# Hot restart (graceful reload)
curl -X POST http://localhost:9901/drain_listeners

# Modify runtime settings
curl -X POST http://localhost:9901/runtime_modify?key1=value1

# View memory usage
curl http://localhost:9901/memory

TLS Configuration

Secure communication with upstream and downstream TLS.

Downstream TLS (Client → Envoy)

listeners:
- name: https_listener
  address:
    socket_address:
      address: 0.0.0.0
      port_value: 443
  filter_chains:
  - filters:
    - name: envoy.filters.network.http_connection_manager
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
        stat_prefix: ingress_https
        route_config:
          name: local_route
          virtual_hosts:
          - name: backend
            domains: ["*"]
            routes:
            - match:
                prefix: "/"
              route:
                cluster: backend_cluster
        http_filters:
        - name: envoy.filters.http.router
          typed_config:
            "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

    # TLS context for this filter chain
    transport_socket:
      name: envoy.transport_sockets.tls
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext
        common_tls_context:
          tls_certificates:
          # Certificate from file
          - certificate_chain:
              filename: "/etc/envoy/certs/server-cert.pem"
            private_key:
              filename: "/etc/envoy/certs/server-key.pem"

          # Validation context (for mTLS)
          validation_context:
            trusted_ca:
              filename: "/etc/envoy/certs/ca-cert.pem"

          # TLS parameters
          tls_params:
            tls_minimum_protocol_version: TLSv1_2
            tls_maximum_protocol_version: TLSv1_3
            cipher_suites:
            - "ECDHE-ECDSA-AES128-GCM-SHA256"
            - "ECDHE-RSA-AES128-GCM-SHA256"
            - "ECDHE-ECDSA-AES256-GCM-SHA384"
            - "ECDHE-RSA-AES256-GCM-SHA384"

          # ALPN protocols
          alpn_protocols:
          - "h2"
          - "http/1.1"

        # Require client certificates (mTLS)
        require_client_certificate: true

Upstream TLS (Envoy → Backend)

clusters:
- name: secure_backend
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  load_assignment:
    cluster_name: secure_backend
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: secure-backend.local
              port_value: 443

  # Upstream TLS context
  transport_socket:
    name: envoy.transport_sockets.tls
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.UpstreamTlsContext
      common_tls_context:
        # Client certificate (for mTLS)
        tls_certificates:
        - certificate_chain:
            filename: "/etc/envoy/certs/client-cert.pem"
          private_key:
            filename: "/etc/envoy/certs/client-key.pem"

        # Server validation
        validation_context:
          trusted_ca:
            filename: "/etc/envoy/certs/ca-cert.pem"
          match_typed_subject_alt_names:
          - san_type: DNS
            matcher:
              exact: "secure-backend.local"

        # ALPN
        alpn_protocols:
        - "h2"

      # SNI to send
      sni: "secure-backend.local"

Certificate Rotation with SDS

# Secret Discovery Service for dynamic cert loading
listeners:
- name: https_listener
  address:
    socket_address:
      address: 0.0.0.0
      port_value: 443
  filter_chains:
  - filters:
    - name: envoy.filters.network.http_connection_manager
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
        stat_prefix: ingress_https
        route_config:
          name: local_route
          virtual_hosts:
          - name: backend
            domains: ["*"]
            routes:
            - match:
                prefix: "/"
              route:
                cluster: backend_cluster
        http_filters:
        - name: envoy.filters.http.router
          typed_config:
            "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

    transport_socket:
      name: envoy.transport_sockets.tls
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext
        common_tls_context:
          # Reference secret via SDS
          tls_certificate_sds_secret_configs:
          - name: server_cert
            sds_config:
              resource_api_version: V3
              api_config_source:
                api_type: GRPC
                transport_api_version: V3
                grpc_services:
                - envoy_grpc:
                    cluster_name: sds_cluster

clusters:
- name: sds_cluster
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  http2_protocol_options: {}
  load_assignment:
    cluster_name: sds_cluster
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: sds-server.local
              port_value: 8234

Quick Reference

Essential Configuration Hierarchy

static_resources:
  listeners:          # Entry points (bind to ports)
    - filter_chains:  # Network/HTTP filter pipelines
        - filters:    # Process connections/requests

  clusters:           # Upstream service groups
    - load_assignment:  # Endpoint configuration
    - health_checks:    # Active health checking
    - outlier_detection: # Passive health checking

admin:                # Admin interface config
stats_sinks:          # Metrics exporters
tracing:              # Distributed tracing

Common Filter Order

http_filters:
- envoy.filters.http.health_check     # Health check responder
- envoy.filters.http.cors             # CORS handling
- envoy.filters.http.ext_authz        # External authorisation
- envoy.filters.http.jwt_authn        # JWT validation
- envoy.filters.http.ratelimit        # Rate limiting
- envoy.filters.http.fault            # Fault injection (testing)
- envoy.filters.http.buffer           # Request buffering
- envoy.filters.http.grpc_web         # gRPC-Web support
- envoy.filters.http.router           # Routing (must be last)

Load Balancing Algorithms

Policy Use Case Characteristics
ROUND_ROBIN General purpose Distributes evenly across endpoints
LEAST_REQUEST Variable request duration Routes to endpoint with fewest active requests
RANDOM Simple distribution Random selection, good for large clusters
RING_HASH Sticky sessions Consistent hashing based on request properties
MAGLEV Sticky sessions Faster consistent hashing, better distribution

Cluster Discovery Types

Type When to Use Refresh Method
STATIC Fixed endpoints, rarely change Manual configuration update
STRICT_DNS DNS with multiple A/AAAA records Periodic DNS resolution
LOGICAL_DNS Single endpoint with dynamic IP DNS lookup per connection
EDS Dynamic service discovery xDS API updates

Retry Conditions

Condition Trigger
5xx Any 5xx response
gateway-error 502, 503, 504 only
reset Connection reset before response
connect-failure Connection failure to upstream
refused-stream HTTP/2 REFUSED_STREAM
retriable-4xx 409 only
retriable-status-codes Custom status codes
retriable-headers Response header indicates retry

Response Flags

Flag Meaning
UH Upstream host connection failure
UF Upstream connection failure
UO Upstream overflow (circuit breaker open)
NR No route configured
URX Request exceeded max retries
NC No healthy upstream
DT Downstream request timeout
UT Upstream request timeout
LR Local rate limited
UR Upstream rate limited
UC Upstream connection termination
DPE Downstream protocol error

Admin API Endpoints

# Configuration
/config_dump              # Full configuration dump
/config_dump?resource=   # Specific resource (listeners, clusters, routes)

# Statistics
/stats                    # All stats (text)
/stats/prometheus         # Prometheus format
/stats?format=json        # JSON format
/stats?filter=cluster     # Filtered stats

# Runtime
/clusters                 # Cluster status and health
/listeners                # Active listeners
/server_info              # Version and uptime
/ready                    # Readiness check
/runtime                  # Runtime values
/runtime_modify?key=val   # Modify runtime setting

# Operations
/drain_listeners?graceful  # Graceful shutdown
/drain_listeners?inboundonly  # Drain inbound only
/healthcheck/fail         # Force health check fail
/healthcheck/ok           # Force health check ok
/logging?level=debug      # Change log level
/reset_counters           # Reset statistics

Environment Variables

# Common Envoy environment variables
ENVOY_UID=0                          # User ID
ENVOY_GID=0                          # Group ID
ENVOY_LOG_LEVEL=info                 # Log level (trace, debug, info, warning, error, critical)
ENVOY_LOG_FORMAT="[%Y-%m-%d %T.%e][%t][%l][%n] %v"  # Log format
ENVOY_COMPONENT_LOG_LEVEL=upstream:debug,connection:trace  # Per-component logging

Common CLI Arguments

# Start Envoy with configuration
envoy -c /etc/envoy/envoy.yaml

# Configuration validation only
envoy --mode validate -c /etc/envoy/envoy.yaml

# Override log level
envoy -c /etc/envoy/envoy.yaml -l debug

# Set component log levels
envoy -c /etc/envoy/envoy.yaml --component-log-level upstream:debug,connection:trace

# Drain time for graceful shutdown
envoy -c /etc/envoy/envoy.yaml --drain-time-s 60

# Service cluster and node for identification
envoy -c /etc/envoy/envoy.yaml \
  --service-cluster my-cluster \
  --service-node my-node-1

# Concurrency (worker threads)
envoy -c /etc/envoy/envoy.yaml --concurrency 4

Common Issues and Solutions

Issue: Upstream Connection Failures

Symptoms: 503 responses, UH or UF response flags in logs.

Solutions:

# Increase connection timeout
clusters:
- name: backend_cluster
  connect_timeout: 5s  # Increase from default 2s

# Verify DNS resolution
- name: backend_cluster
  type: STRICT_DNS
  dns_lookup_family: V4_ONLY  # or V6_ONLY, AUTO

# Add health checks to detect failed endpoints
  health_checks:
  - timeout: 1s
    interval: 5s
    unhealthy_threshold: 2
    healthy_threshold: 2
    http_health_check:
      path: "/health"

Debugging:

# Check cluster status
curl http://localhost:9901/clusters | grep backend_cluster

# Look for connection errors in logs
grep "upstream connect error" /var/log/envoy/envoy.log

Issue: Circuit Breaker Tripping

Symptoms: 503 responses with UO (upstream overflow) flag.

Solutions:

# Increase circuit breaker thresholds
clusters:
- name: backend_cluster
  circuit_breakers:
    thresholds:
    - priority: DEFAULT
      max_connections: 2048       # Increase from 1024
      max_pending_requests: 2048
      max_requests: 2048
      max_retries: 5

# Check circuit breaker stats
curl http://localhost:9901/stats | grep circuit_breakers

Prevention:

# Add outlier detection to remove slow endpoints
outlier_detection:
  consecutive_5xx: 5
  interval: 10s
  base_ejection_time: 30s
  max_ejection_percent: 50

Issue: High Latency

Symptoms: Slow response times, high DURATION values in access logs.

Solutions:

# Optimise timeouts
routes:
- match:
    prefix: "/api"
  route:
    cluster: backend_cluster
    timeout: 10s           # Reduce total timeout
    idle_timeout: 300s
    per_try_timeout: 3s    # Set per-retry timeout

# Use LEAST_REQUEST load balancing for variable latency
clusters:
- name: backend_cluster
  lb_policy: LEAST_REQUEST
  least_request_lb_config:
    choice_count: 2  # Compare 2 endpoints per request

# Enable HTTP/2 connection pooling
  http2_protocol_options:
    max_concurrent_streams: 100

Debugging:

# Check upstream service time
curl http://localhost:9901/stats | grep upstream_rq_time

# Review access logs for slow requests
grep -E '"duration":[0-9]{4,}' /var/log/envoy/access.log

Issue: TLS Handshake Failures

Symptoms: Connection failures with TLS errors in logs.

Solutions:

# Verify certificate paths and permissions
ls -la /etc/envoy/certs/
# Ensure Envoy process can read certificate files

# Check TLS version compatibility
common_tls_context:
  tls_params:
    tls_minimum_protocol_version: TLSv1_2  # Lower if needed
    tls_maximum_protocol_version: TLSv1_3

# Verify CA trust chain
validation_context:
  trusted_ca:
    filename: "/etc/envoy/certs/ca-bundle.pem"  # Include full chain

# Disable certificate validation (testing only!)
validation_context:
  trust_chain_verification: ACCEPT_UNTRUSTED

Debugging:

# Test TLS connection manually
openssl s_client -connect backend.local:443 -servername backend.local

# Check Envoy TLS stats
curl http://localhost:9901/stats | grep ssl

Issue: Memory Leaks

Symptoms: Gradually increasing memory usage, eventual OOM kills.

Solutions:

# Limit connection buffer sizes
per_connection_buffer_limit_bytes: 32768  # Default 1MB

# Configure stats flush interval
stats_flush_interval: 5s  # Default 5s

# Limit access log buffer
access_log:
- name: envoy.access_loggers.file
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
    path: "/var/log/envoy/access.log"

Monitoring:

# Check memory usage
curl http://localhost:9901/memory

# Monitor heap allocation
curl http://localhost:9901/stats | grep memory

Issue: Configuration Validation Errors

Symptoms: Envoy fails to start with validation errors.

Solutions:

# Validate configuration before deployment
envoy --mode validate -c /etc/envoy/envoy.yaml

# Check for common issues:
# 1. Missing required fields
# 2. Incorrect @type in typed_config
# 3. Invalid resource references (cluster not defined)
# 4. YAML syntax errors

# Use strict validation
envoy --mode validate -c /etc/envoy/envoy.yaml --reject-unknown-dynamic-fields

Issue: xDS Configuration Update Failures

Symptoms: Configuration not updating, stale routes/clusters.

Solutions:

# Add timeout for xDS connections
dynamic_resources:
  cds_config:
    resource_api_version: V3
    api_config_source:
      api_type: GRPC
      transport_api_version: V3
      grpc_services:
      - envoy_grpc:
          cluster_name: xds_cluster
      set_node_on_first_message_only: true
      request_timeout: 10s  # Add timeout

# Enable xDS debug logging
--component-log-level upstream:debug,config:debug

Debugging:

# Check xDS connection status
curl http://localhost:9901/config_dump | jq '.configs[].dynamic_active_clusters'

# Monitor xDS stats
curl http://localhost:9901/stats | grep control_plane

Issue: Rate Limiting Not Working

Symptoms: Rate limits not enforced, excessive requests reaching upstream.

Solutions:

# Verify filter order (ratelimit before router)
http_filters:
- name: envoy.filters.http.local_ratelimit
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.local_ratelimit.v3.LocalRateLimit
    stat_prefix: http_local_rate_limiter
    token_bucket:
      max_tokens: 100
      tokens_per_fill: 10
      fill_interval: 1s
    filter_enabled:
      default_value:
        numerator: 100
        denominator: HUNDRED
    filter_enforced:
      default_value:
        numerator: 100    # Ensure enforced
        denominator: HUNDRED

- name: envoy.filters.http.router
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

Debugging:

# Check rate limit stats
curl http://localhost:9901/stats | grep local_rate_limit

# Look for rate limit response headers
curl -I http://localhost:8080/api

Issue: Health Check Failures

Symptoms: All endpoints marked unhealthy, no traffic forwarded.

Solutions:

# Adjust health check thresholds
health_checks:
- timeout: 2s             # Increase timeout
  interval: 10s           # Reduce frequency
  unhealthy_threshold: 5  # Allow more failures
  healthy_threshold: 2
  http_health_check:
    path: "/health"
    expected_statuses:
    - start: 200
      end: 399  # Accept 3xx responses

# Add service_name_matcher for HTTP/2
  http_health_check:
    path: "/health"
    host: "backend.local"  # Add host header

Debugging:

# Check health check stats per endpoint
curl http://localhost:9901/clusters | grep health

# Test health endpoint manually
curl http://backend.local:8080/health

Issue: gRPC Services Not Working

Symptoms: gRPC clients cannot connect, errors about HTTP/2.

Solutions:

# Ensure HTTP/2 support enabled
http_connection_manager:
  codec_type: AUTO  # or HTTP2
  http2_protocol_options:
    initial_stream_window_size: 65536
    initial_connection_window_size: 1048576

  http_filters:
  # Add gRPC-Web filter if needed
  - name: envoy.filters.http.grpc_web
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.http.grpc_web.v3.GrpcWeb

  - name: envoy.filters.http.router
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

# Configure upstream for gRPC
clusters:
- name: grpc_backend
  type: STRICT_DNS
  lb_policy: ROUND_ROBIN
  http2_protocol_options: {}  # Enable HTTP/2
  load_assignment:
    cluster_name: grpc_backend
    endpoints:
    - lb_endpoints:
      - endpoint:
          address:
            socket_address:
              address: grpc-service.local
              port_value: 9090

Issue: Headers Too Large

Symptoms: 431 Request Header Fields Too Large errors.

Solutions:

# Increase header limits
http_connection_manager:
  stat_prefix: ingress_http
  codec_type: AUTO
  # max_request_headers_kb is a field of the HCM itself, not common_http_protocol_options
  max_request_headers_kb: 96        # Default 60KB
  common_http_protocol_options:
    max_headers_count: 100          # Default 100
  http2_protocol_options:
    max_concurrent_streams: 100

Complete Production Example

# Production-ready Envoy configuration
admin:
  address:
    socket_address:
      address: 127.0.0.1
      port_value: 9901

static_resources:
  listeners:
  - name: https_listener
    address:
      socket_address:
        address: 0.0.0.0
        port_value: 443

    per_connection_buffer_limit_bytes: 32768

    filter_chains:
    - filters:
      - name: envoy.filters.network.http_connection_manager
        typed_config:
          "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
          stat_prefix: ingress_https
          codec_type: AUTO
          use_remote_address: true
          xff_num_trusted_hops: 1

          max_request_headers_kb: 96

          common_http_protocol_options:
            idle_timeout: 300s

          http2_protocol_options:
            max_concurrent_streams: 100

          route_config:
            name: local_route
            virtual_hosts:
            - name: api_service
              domains: ["api.example.com"]
              routes:
              - match:
                  prefix: "/health"
                direct_response:
                  status: 200
                  body:
                    inline_string: "OK"

              - match:
                  prefix: "/api/v1"
                route:
                  cluster: api_v1_cluster
                  timeout: 15s
                  idle_timeout: 60s
                  retry_policy:
                    retry_on: "5xx,reset,connect-failure"
                    num_retries: 3
                    per_try_timeout: 5s
                    retry_back_off:
                      base_interval: 0.1s
                      max_interval: 1s

          http_filters:
          - name: envoy.filters.http.cors
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.cors.v3.Cors

          - name: envoy.filters.http.ext_authz
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthz
              grpc_service:
                envoy_grpc:
                  cluster_name: auth_service
                timeout: 0.5s
              failure_mode_allow: false

          - name: envoy.filters.http.local_ratelimit
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.local_ratelimit.v3.LocalRateLimit
              stat_prefix: http_local_rate_limiter
              token_bucket:
                max_tokens: 1000
                tokens_per_fill: 100
                fill_interval: 1s
              filter_enabled:
                default_value:
                  numerator: 100
                  denominator: HUNDRED
              filter_enforced:
                default_value:
                  numerator: 100
                  denominator: HUNDRED

          - name: envoy.filters.http.router
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

          access_log:
          - name: envoy.access_loggers.file
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
              path: "/var/log/envoy/access.log"
              json_format:
                timestamp: "%START_TIME%"
                method: "%REQ(:METHOD)%"
                path: "%REQ(X-ENVOY-ORIGINAL-PATH?:PATH)%"
                protocol: "%PROTOCOL%"
                response_code: "%RESPONSE_CODE%"
                response_flags: "%RESPONSE_FLAGS%"
                bytes_received: "%BYTES_RECEIVED%"
                bytes_sent: "%BYTES_SENT%"
                duration: "%DURATION%"
                upstream_service_time: "%RESP(X-ENVOY-UPSTREAM-SERVICE-TIME)%"
                user_agent: "%REQ(USER-AGENT)%"
                request_id: "%REQ(X-REQUEST-ID)%"
                upstream_host: "%UPSTREAM_HOST%"

      transport_socket:
        name: envoy.transport_sockets.tls
        typed_config:
          "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext
          common_tls_context:
            tls_certificates:
            - certificate_chain:
                filename: "/etc/envoy/certs/server-cert.pem"
              private_key:
                filename: "/etc/envoy/certs/server-key.pem"
            tls_params:
              tls_minimum_protocol_version: TLSv1_2
              tls_maximum_protocol_version: TLSv1_3
              cipher_suites:
              - "ECDHE-ECDSA-AES128-GCM-SHA256"
              - "ECDHE-RSA-AES128-GCM-SHA256"
            alpn_protocols:
            - "h2"
            - "http/1.1"

  clusters:
  - name: api_v1_cluster
    type: STRICT_DNS
    lb_policy: LEAST_REQUEST
    connect_timeout: 2s

    circuit_breakers:
      thresholds:
      - priority: DEFAULT
        max_connections: 1024
        max_pending_requests: 1024
        max_requests: 1024
        max_retries: 3

    outlier_detection:
      consecutive_5xx: 5
      interval: 10s
      base_ejection_time: 30s
      max_ejection_percent: 50

    health_checks:
    - timeout: 1s
      interval: 10s
      unhealthy_threshold: 3
      healthy_threshold: 2
      http_health_check:
        path: "/health"

    load_assignment:
      cluster_name: api_v1_cluster
      endpoints:
      - lb_endpoints:
        - endpoint:
            address:
              socket_address:
                address: api-v1.local
                port_value: 8080

    transport_socket:
      name: envoy.transport_sockets.tls
      typed_config:
        "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.UpstreamTlsContext
        common_tls_context:
          validation_context:
            trusted_ca:
              filename: "/etc/envoy/certs/ca-cert.pem"
        sni: "api-v1.local"

  - name: auth_service
    type: STRICT_DNS
    lb_policy: ROUND_ROBIN
    connect_timeout: 1s
    http2_protocol_options: {}
    load_assignment:
      cluster_name: auth_service
      endpoints:
      - lb_endpoints:
        - endpoint:
            address:
              socket_address:
                address: auth.local
                port_value: 9191

  # statsd_cluster is referenced by the stats_sink below; it must be a
  # member of static_resources.clusters, not a top-level key.
  - name: statsd_cluster
    type: STRICT_DNS
    lb_policy: ROUND_ROBIN
    connect_timeout: 1s
    load_assignment:
      cluster_name: statsd_cluster
      endpoints:
      - lb_endpoints:
        - endpoint:
            address:
              socket_address:
                address: statsd.local
                port_value: 8125

stats_sinks:
- name: envoy.stat_sinks.statsd
  typed_config:
    "@type": type.googleapis.com/envoy.config.metrics.v3.StatsdSink
    tcp_cluster_name: statsd_cluster
    prefix: "envoy"