Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Jaeger

An open-source distributed tracing system for monitoring and troubleshooting microservices-based architectures.

Jaeger Cheatsheet

An open-source distributed tracing system for monitoring and troubleshooting microservices-based architectures.


Version note (important): This cheatsheet documents Jaeger v1, which reached end-of-life on 31 December 2025. The current release is Jaeger v2 (built on the OpenTelemetry Collector), which ships a single jaegertracing/jaeger binary replacing the separate v1 all-in-one/collector/query/ingester images and the Jaeger Agent. The jaeger-client instrumentation libraries shown below were retired in 2022 in favour of the OpenTelemetry SDKs — use OpenTelemetry for all new instrumentation (see the Integration with OpenTelemetry section). The v1 material is retained for existing deployments; migrate to v2 + OTel for anything new.


Overview

Jaeger is a distributed tracing platform originally developed by Uber and now a CNCF graduated project. It helps developers monitor and troubleshoot transactions in complex distributed systems by visualising request flows across services, identifying performance bottlenecks, and analysing root causes of errors.

Storage LayerJaeger ArchitectureApplication LayerService AService BService CJaeger AgentJaeger CollectorJaeger QueryJaeger UIStorage BackendStorage LayerJaeger ArchitectureApplication LayerService AService BService CJaeger AgentJaeger CollectorJaeger QueryJaeger UIStorage Backend

Core Features

Feature Description
Distributed Context Propagation Tracks requests across service boundaries
Distributed Transaction Monitoring Monitors end-to-end request flows
Root Cause Analysis Identifies performance bottlenecks and errors
Service Dependency Analysis Visualises service relationships
Performance Optimisation Measures latency and throughput

Architecture

Jaeger's architecture consists of several components that work together to collect, store, and visualise traces.

PresentationStorageData ProcessingData CollectionThrift/UDPgRPCInstrumented AppAgentUDP 6831/6832CollectorgRPC 14250HTTP 14268KafkaOptionalIngesterElasticsearchCassandraMemoryQuery ServiceHTTP 16686Web UIPresentationStorageData ProcessingData CollectionThrift/UDPgRPCInstrumented AppAgentUDP 6831/6832CollectorgRPC 14250HTTP 14268KafkaOptionalIngesterElasticsearchCassandraMemoryQuery ServiceHTTP 16686Web UI

Jaeger Agent

The agent is a daemon that listens for spans sent over UDP, batches them, and forwards to the collector. It's typically deployed as a sidecar or DaemonSet.

# Kubernetes DaemonSet deployment
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: jaeger-agent
spec:
  selector:
    matchLabels:
      app: jaeger-agent
  template:
    metadata:
      labels:
        app: jaeger-agent
    spec:
      containers:
        - name: jaeger-agent
          image: jaegertracing/jaeger-agent:1.52
          ports:
            - containerPort: 6831
              protocol: UDP
              name: compact-thrift
            - containerPort: 6832
              protocol: UDP
              name: binary-thrift
            - containerPort: 5778
              protocol: TCP
              name: config
          args:
            - --reporter.grpc.host-port=jaeger-collector:14250
            - --log-level=info

Jaeger Collector

The collector receives traces from agents, validates them, and stores them in a backend storage system.

# Docker Compose collector configuration
jaeger-collector:
  image: jaegertracing/jaeger-collector:1.52
  ports:
    - "14269:14269"   # Admin health check
    - "14268:14268"   # HTTP collector
    - "14250:14250"   # gRPC collector
    - "4317:4317"     # OTLP gRPC
    - "4318:4318"     # OTLP HTTP
  environment:
    - SPAN_STORAGE_TYPE=elasticsearch
    - ES_SERVER_URLS=http://elasticsearch:9200
    - ES_TAGS_AS_FIELDS_ALL=true
  command:
    - --collector.queue-size=2000
    - --collector.num-workers=50

Jaeger Query Service

The query service exposes APIs to retrieve traces and hosts the web UI.

# Docker Compose query service
jaeger-query:
  image: jaegertracing/jaeger-query:1.52
  ports:
    - "16686:16686"   # Web UI
    - "16687:16687"   # Admin health check
  environment:
    - SPAN_STORAGE_TYPE=elasticsearch
    - ES_SERVER_URLS=http://elasticsearch:9200
  depends_on:
    - elasticsearch

All-in-One Deployment

For development and testing, use the all-in-one image:

# Run Jaeger all-in-one with in-memory storage
# Note: on Jaeger v1 before 1.46, OTLP ingestion (ports 4317/4318) is
# OFF by default and must be enabled with COLLECTOR_OTLP_ENABLED=true,
# otherwise those ports silently accept nothing. From 1.46 it defaults on.
# The env var is harmless on newer versions so including it aids portability.
docker run -d --name jaeger \
  -e COLLECTOR_OTLP_ENABLED=true \
  -p 6831:6831/udp \
  -p 6832:6832/udp \
  -p 5778:5778 \
  -p 16686:16686 \
  -p 4317:4317 \
  -p 4318:4318 \
  -p 14250:14250 \
  -p 14268:14268 \
  -p 14269:14269 \
  jaegertracing/all-in-one:1.52

# Access UI at http://localhost:16686

Port Reference

Port Protocol Component Purpose
6831 UDP Agent Thrift compact protocol
6832 UDP Agent Thrift binary protocol
5778 HTTP Agent Configuration/sampling
14250 gRPC Collector Model proto
14268 HTTP Collector Jaeger thrift
14269 HTTP Collector Admin/health check
4317 gRPC Collector OTLP gRPC
4318 HTTP Collector OTLP HTTP
16686 HTTP Query Web UI
16687 HTTP Query Admin/health check

Instrumentation Setup

Key Concepts

Span: A logical unit of work with operation name, start time, duration, and tags.

Trace: A collection of spans forming a directed acyclic graph (DAG).

SpanContext: Carries trace identity across process boundaries.

Tags: Key-value pairs providing metadata about a span.

Logs: Timestamped events within a span.

Python Instrumentation

# Install dependencies
# pip install jaeger-client opentracing

from jaeger_client import Config
from opentracing.ext import tags
from opentracing.propagation import Format

# Initialise tracer
def init_tracer(service_name):
    config = Config(
        config={
            'sampler': {
                'type': 'const',
                'param': 1,  # Sample all traces
            },
            'local_agent': {
                'reporting_host': 'localhost',
                'reporting_port': '6831',
            },
            'logging': True,
        },
        service_name=service_name,
        validate=True,
    )
    return config.initialize_tracer()

tracer = init_tracer('my-service')

# Create spans
def process_request(request):
    with tracer.start_active_span('process-request') as scope:
        span = scope.span
        span.set_tag(tags.HTTP_METHOD, 'GET')
        span.set_tag(tags.HTTP_URL, request.url)
        span.set_tag('user.id', request.user_id)

        # Log events within span
        span.log_kv({'event': 'processing started'})

        result = do_work()

        span.log_kv({'event': 'processing completed', 'result': result})
        span.set_tag(tags.HTTP_STATUS_CODE, 200)

        return result

# Propagate context to downstream services
def call_downstream_service(url):
    with tracer.start_active_span('call-downstream') as scope:
        span = scope.span

        # Inject trace context into headers
        headers = {}
        tracer.inject(span.context, Format.HTTP_HEADERS, headers)

        response = requests.get(url, headers=headers)
        span.set_tag(tags.HTTP_STATUS_CODE, response.status_code)

        return response

# Extract context from incoming request
def handle_incoming_request(request):
    # Extract parent span context
    span_context = tracer.extract(Format.HTTP_HEADERS, request.headers)

    with tracer.start_active_span('handle-request', child_of=span_context) as scope:
        # Process request
        return process(request)

Go Instrumentation

package main

import (
    "context"
    "io"
    "net/http"

    "github.com/opentracing/opentracing-go"
    "github.com/opentracing/opentracing-go/ext"
    "github.com/uber/jaeger-client-go"
    "github.com/uber/jaeger-client-go/config"
)

// Initialise tracer
func initTracer(serviceName string) (opentracing.Tracer, io.Closer, error) {
    cfg := config.Configuration{
        ServiceName: serviceName,
        Sampler: &config.SamplerConfig{
            Type:  jaeger.SamplerTypeConst,
            Param: 1,
        },
        Reporter: &config.ReporterConfig{
            LocalAgentHostPort: "localhost:6831",
            LogSpans:           true,
        },
    }

    tracer, closer, err := cfg.NewTracer()
    if err != nil {
        return nil, nil, err
    }

    opentracing.SetGlobalTracer(tracer)
    return tracer, closer, nil
}

func main() {
    tracer, closer, err := initTracer("my-service")
    if err != nil {
        panic(err)
    }
    defer closer.Close()

    // Create root span
    span := tracer.StartSpan("main-operation")
    defer span.Finish()

    ctx := opentracing.ContextWithSpan(context.Background(), span)
    processData(ctx)
}

func processData(ctx context.Context) {
    // Create child span
    span, ctx := opentracing.StartSpanFromContext(ctx, "process-data")
    defer span.Finish()

    span.SetTag("data.type", "user")
    span.LogKV("event", "processing started", "items", 100)

    // HTTP client with tracing
    callService(ctx, "http://service-b/api/data")
}

func callService(ctx context.Context, url string) {
    span, _ := opentracing.StartSpanFromContext(ctx, "http-call")
    defer span.Finish()

    ext.SpanKindRPCClient.Set(span)
    ext.HTTPUrl.Set(span, url)
    ext.HTTPMethod.Set(span, "GET")

    req, _ := http.NewRequest("GET", url, nil)

    // Inject trace context
    opentracing.GlobalTracer().Inject(
        span.Context(),
        opentracing.HTTPHeaders,
        opentracing.HTTPHeadersCarrier(req.Header),
    )

    client := &http.Client{}
    resp, err := client.Do(req)
    if err != nil {
        ext.LogError(span, err)
        return
    }
    defer resp.Body.Close()

    ext.HTTPStatusCode.Set(span, uint16(resp.StatusCode))
}

Java Instrumentation

import io.jaegertracing.Configuration;
import io.jaegertracing.Configuration.SamplerConfiguration;
import io.jaegertracing.Configuration.ReporterConfiguration;
import io.opentracing.Scope;
import io.opentracing.Span;
import io.opentracing.Tracer;
import io.opentracing.propagation.Format;
import io.opentracing.propagation.TextMapAdapter;
import io.opentracing.tag.Tags;

public class TracingExample {

    private static Tracer initTracer(String serviceName) {
        SamplerConfiguration sampler = SamplerConfiguration.fromEnv()
            .withType("const")
            .withParam(1);

        ReporterConfiguration reporter = ReporterConfiguration.fromEnv()
            .withLogSpans(true);

        return new Configuration(serviceName)
            .withSampler(sampler)
            .withReporter(reporter)
            .getTracer();
    }

    public static void main(String[] args) {
        Tracer tracer = initTracer("my-service");

        // Create span with try-with-resources
        try (Scope scope = tracer.buildSpan("main-operation")
                .withTag("environment", "production")
                .startActive(true)) {

            Span span = scope.span();
            span.log("Processing started");

            processRequest(tracer);

            span.log("Processing completed");
        }

        tracer.close();
    }

    private static void processRequest(Tracer tracer) {
        Span parentSpan = tracer.activeSpan();

        try (Scope scope = tracer.buildSpan("process-request")
                .asChildOf(parentSpan)
                .withTag(Tags.SPAN_KIND.getKey(), Tags.SPAN_KIND_SERVER)
                .startActive(true)) {

            Span span = scope.span();
            span.setTag("request.type", "api");

            // Inject context for downstream call
            Map<String, String> headers = new HashMap<>();
            tracer.inject(
                span.context(),
                Format.Builtin.HTTP_HEADERS,
                new TextMapAdapter(headers)
            );

            // Make HTTP call with headers
            callDownstream(headers);
        }
    }
}

Node.js Instrumentation

const { initTracer } = require('jaeger-client');
const opentracing = require('opentracing');

// Initialise tracer
const config = {
    serviceName: 'my-service',
    sampler: {
        type: 'const',
        param: 1,
    },
    reporter: {
        agentHost: 'localhost',
        agentPort: 6831,
        logSpans: true,
    },
};

const options = {
    tags: {
        'my-service.version': '1.0.0',
    },
};

const tracer = initTracer(config, options);

// Create spans
async function processRequest(req, res) {
    // Extract context from incoming request
    const parentSpanContext = tracer.extract(
        opentracing.FORMAT_HTTP_HEADERS,
        req.headers
    );

    const span = tracer.startSpan('process-request', {
        childOf: parentSpanContext,
        tags: {
            [opentracing.Tags.SPAN_KIND]: opentracing.Tags.SPAN_KIND_RPC_SERVER,
            [opentracing.Tags.HTTP_METHOD]: req.method,
            [opentracing.Tags.HTTP_URL]: req.url,
        },
    });

    try {
        span.log({ event: 'processing started' });

        const result = await doWork();

        span.setTag(opentracing.Tags.HTTP_STATUS_CODE, 200);
        span.log({ event: 'processing completed', result });

        return result;
    } catch (error) {
        span.setTag(opentracing.Tags.ERROR, true);
        span.log({
            event: 'error',
            'error.object': error,
            message: error.message,
            stack: error.stack,
        });
        throw error;
    } finally {
        span.finish();
    }
}

// Propagate context to downstream service
async function callDownstream(parentSpan, url) {
    const span = tracer.startSpan('call-downstream', {
        childOf: parentSpan,
        tags: {
            [opentracing.Tags.SPAN_KIND]: opentracing.Tags.SPAN_KIND_RPC_CLIENT,
            [opentracing.Tags.HTTP_URL]: url,
        },
    });

    const headers = {};
    tracer.inject(span.context(), opentracing.FORMAT_HTTP_HEADERS, headers);

    try {
        const response = await fetch(url, { headers });
        span.setTag(opentracing.Tags.HTTP_STATUS_CODE, response.status);
        return response;
    } finally {
        span.finish();
    }
}

Trace Viewing and Analysis

Using the Jaeger UI

The Jaeger UI provides powerful visualisation and analysis capabilities:

Trace DetailResults ViewSearch PanelSelect ServiceSelect OperationFilter by TagsTime RangeResult LimitTrace ListTimeline ViewService GraphWaterfall ViewSpan DetailsSpan LogsCompare TracesTrace DetailResults ViewSearch PanelSelect ServiceSelect OperationFilter by TagsTime RangeResult LimitTrace ListTimeline ViewService GraphWaterfall ViewSpan DetailsSpan LogsCompare Traces

Searching Traces

# Search by service name
# UI: Select service from dropdown

# Search by operation name
# UI: Select operation after choosing service

# Search by tags (key=value)
http.status_code=500
user.id=12345
error=true

# Search by duration
# minDuration: 100ms
# maxDuration: 5s

# Search by trace ID
# Direct lookup: http://localhost:16686/trace/<trace-id>

Common Tag Queries

Query Purpose
error=true Find all error traces
http.status_code=500 Find server errors
http.status_code=404 Find not found errors
http.method=POST Find POST requests
db.type=postgresql Find database operations
kafka.topic=orders Find Kafka messages

Trace Comparison

Compare two traces to identify differences:

  1. Select first trace and click "Compare"
  2. Select second trace
  3. View side-by-side comparison
  4. Identify timing differences and missing spans

API Queries

# Get all services
curl "http://localhost:16686/api/services"

# Get operations for a service
curl "http://localhost:16686/api/services/my-service/operations"

# Search traces
curl "http://localhost:16686/api/traces?service=my-service&limit=20"

# Get trace by ID
curl "http://localhost:16686/api/traces/<trace-id>"

# Search with tags
curl "http://localhost:16686/api/traces?service=my-service&tags=%7B%22error%22%3A%22true%22%7D"

# Get dependencies
curl "http://localhost:16686/api/dependencies?endTs=$(date +%s)000&lookback=3600000"

Analysing Performance

# Python script to analyse trace data
import requests
import statistics

def analyse_traces(service_name, operation, limit=100):
    """Analyse trace latencies for a service operation."""
    url = f"http://localhost:16686/api/traces"
    params = {
        "service": service_name,
        "operation": operation,
        "limit": limit
    }

    response = requests.get(url, params=params)
    traces = response.json()['data']

    durations = []
    for trace in traces:
        # Get root span duration
        for span in trace['spans']:
            if span['operationName'] == operation:
                duration_us = span['duration']
                durations.append(duration_us / 1000)  # Convert to ms
                break

    if durations:
        print(f"Traces analysed: {len(durations)}")
        print(f"Min latency: {min(durations):.2f}ms")
        print(f"Max latency: {max(durations):.2f}ms")
        print(f"Mean latency: {statistics.mean(durations):.2f}ms")
        print(f"Median latency: {statistics.median(durations):.2f}ms")
        print(f"P95 latency: {sorted(durations)[int(len(durations)*0.95)]:.2f}ms")
        print(f"P99 latency: {sorted(durations)[int(len(durations)*0.99)]:.2f}ms")

analyse_traces("api-gateway", "GET /api/orders")

Sampling Strategies

Key Concepts

Sampling determines which traces to record. Different strategies balance visibility with overhead:

Strategy Description Use Case
Const Sample all (1) or none (0) Development/debugging
Probabilistic Sample percentage of traces Production baseline
Rate Limiting Fixed traces per second Cost control
Remote Server-controlled sampling Dynamic adjustment
Adaptive Automatic adjustment High-traffic systems

Configuration Examples

# Const sampler - sample everything
sampler:
  type: const
  param: 1  # 1 = sample all, 0 = sample none

# Probabilistic sampler - 10% of traces
sampler:
  type: probabilistic
  param: 0.1

# Rate limiting sampler - 2 traces per second
sampler:
  type: ratelimiting
  param: 2

# Remote sampler - fetch from agent
sampler:
  type: remote
  param: 1
  sampling_server_url: http://localhost:5778/sampling

Remote Sampling Configuration

Configure the collector to serve sampling strategies:

// sampling-strategies.json
{
  "service_strategies": [
    {
      "service": "api-gateway",
      "type": "probabilistic",
      "param": 0.5
    },
    {
      "service": "payment-service",
      "type": "probabilistic",
      "param": 1.0
    },
    {
      "service": "logging-service",
      "type": "ratelimiting",
      "param": 10
    }
  ],
  "default_strategy": {
    "type": "probabilistic",
    "param": 0.1
  },
  "per_operation_strategies": [
    {
      "service": "api-gateway",
      "operation_strategies": [
        {
          "operation": "GET /health",
          "type": "probabilistic",
          "param": 0.001
        },
        {
          "operation": "POST /api/orders",
          "type": "probabilistic",
          "param": 1.0
        }
      ]
    }
  ]
}
# Start collector with sampling strategies
docker run -d \
  -p 14268:14268 \
  -p 5778:5778 \
  -v $(pwd)/sampling-strategies.json:/etc/jaeger/sampling-strategies.json \
  jaegertracing/jaeger-collector:1.52 \
  --sampling.strategies-file=/etc/jaeger/sampling-strategies.json

Adaptive Sampling

# Collector configuration for adaptive sampling
collector:
  sampling:
    strategies-file: /etc/jaeger/sampling-strategies.json
    strategies-reload-interval: 30s

# Adaptive sampling (Jaeger 1.27+)
sampling:
  default_strategy:
    type: adaptive
    adaptive:
      sampling_store: memory
      target_samples_per_second: 1.0
      delta_tolerance: 0.3
      buckets_for_calculation: 1
      calculation_interval: 1m

Priority-Based Sampling

// Go - Priority tag for guaranteed sampling
func createHighPrioritySpan(tracer opentracing.Tracer) {
    span := tracer.StartSpan("critical-operation")

    // Set sampling priority to ensure trace is recorded
    ext.SamplingPriority.Set(span, 1)

    defer span.Finish()
}

// In error handling
func handleError(span opentracing.Span, err error) {
    // Force sampling on errors
    ext.SamplingPriority.Set(span, 1)
    ext.LogError(span, err)
}

Storage Backends

Supported Storage Options

Cloud StorageDevelopment StorageProduction StorageJaeger CollectorElasticsearchCassandraKafka + FlinkMemoryBadgerElasticsearchServiceCassandra as aServiceCloud StorageDevelopment StorageProduction StorageJaeger CollectorElasticsearchCassandraKafka + FlinkMemoryBadgerElasticsearchServiceCassandra as aService

Elasticsearch Configuration

# Docker Compose with Elasticsearch
version: '3.8'

services:
  elasticsearch:
    image: docker.elastic.co/elasticsearch/elasticsearch:8.11.0
    environment:
      - discovery.type=single-node
      - xpack.security.enabled=false
      - "ES_JAVA_OPTS=-Xms1g -Xmx1g"
    ports:
      - "9200:9200"
    volumes:
      - esdata:/usr/share/elasticsearch/data

  jaeger-collector:
    image: jaegertracing/jaeger-collector:1.52
    environment:
      - SPAN_STORAGE_TYPE=elasticsearch
      - ES_SERVER_URLS=http://elasticsearch:9200
      - ES_NUM_SHARDS=5
      - ES_NUM_REPLICAS=1
      - ES_INDEX_PREFIX=jaeger
      - ES_TAGS_AS_FIELDS_ALL=true
      - ES_BULK_SIZE=5000000
      - ES_BULK_WORKERS=2
      - ES_BULK_ACTIONS=1000
      - ES_BULK_FLUSH_INTERVAL=200ms
    ports:
      - "14268:14268"
      - "14250:14250"

  jaeger-query:
    image: jaegertracing/jaeger-query:1.52
    environment:
      - SPAN_STORAGE_TYPE=elasticsearch
      - ES_SERVER_URLS=http://elasticsearch:9200
      - ES_INDEX_PREFIX=jaeger
    ports:
      - "16686:16686"

volumes:
  esdata:

Cassandra Configuration

# Docker Compose with Cassandra
version: '3.8'

services:
  cassandra:
    image: cassandra:4.1
    ports:
      - "9042:9042"
    volumes:
      - cassdata:/var/lib/cassandra
    environment:
      - CASSANDRA_DC=dc1
      - CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch

  cassandra-schema:
    image: jaegertracing/jaeger-cassandra-schema:1.52
    depends_on:
      - cassandra
    environment:
      - CQLSH_HOST=cassandra
      - DATACENTER=dc1
      - MODE=prod

  jaeger-collector:
    image: jaegertracing/jaeger-collector:1.52
    environment:
      - SPAN_STORAGE_TYPE=cassandra
      - CASSANDRA_SERVERS=cassandra
      - CASSANDRA_KEYSPACE=jaeger_v1_dc1
    ports:
      - "14268:14268"
      - "14250:14250"
    depends_on:
      - cassandra-schema

  jaeger-query:
    image: jaegertracing/jaeger-query:1.52
    environment:
      - SPAN_STORAGE_TYPE=cassandra
      - CASSANDRA_SERVERS=cassandra
      - CASSANDRA_KEYSPACE=jaeger_v1_dc1
    ports:
      - "16686:16686"

volumes:
  cassdata:

Kafka Streaming Pipeline

# High-throughput setup with Kafka
version: '3.8'

services:
  kafka:
    image: confluentinc/cp-kafka:7.5.0
    environment:
      KAFKA_BROKER_ID: 1
      KAFKA_ZOOKEEPER_CONNECT: zookeeper:2181
      KAFKA_ADVERTISED_LISTENERS: PLAINTEXT://kafka:9092
      KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1

  jaeger-collector:
    image: jaegertracing/jaeger-collector:1.52
    environment:
      - SPAN_STORAGE_TYPE=kafka
      - KAFKA_PRODUCER_BROKERS=kafka:9092
      - KAFKA_PRODUCER_TOPIC=jaeger-spans
    ports:
      - "14268:14268"

  jaeger-ingester:
    image: jaegertracing/jaeger-ingester:1.52
    environment:
      - SPAN_STORAGE_TYPE=elasticsearch
      - KAFKA_CONSUMER_BROKERS=kafka:9092
      - KAFKA_CONSUMER_TOPIC=jaeger-spans
      - KAFKA_CONSUMER_GROUP=jaeger-ingester
      - ES_SERVER_URLS=http://elasticsearch:9200
      - INGESTER_PARALLELISM=1000
      - INGESTER_DEADLOCKINTERVAL=5m

Badger (Local Storage)

# For development/testing with persistent storage
docker run -d \
  --name jaeger \
  -p 16686:16686 \
  -p 6831:6831/udp \
  -v jaeger_data:/badger \
  -e SPAN_STORAGE_TYPE=badger \
  -e BADGER_EPHEMERAL=false \
  -e BADGER_DIRECTORY_VALUE=/badger/data \
  -e BADGER_DIRECTORY_KEY=/badger/key \
  jaegertracing/all-in-one:1.52

Index Management

# Elasticsearch index lifecycle management
# Create ILM policy
curl -X PUT "localhost:9200/_ilm/policy/jaeger-policy" -H 'Content-Type: application/json' -d'
{
  "policy": {
    "phases": {
      "hot": {
        "min_age": "0ms",
        "actions": {
          "rollover": {
            "max_age": "1d",
            "max_size": "50gb"
          }
        }
      },
      "delete": {
        "min_age": "7d",
        "actions": {
          "delete": {}
        }
      }
    }
  }
}'

# Apply to Jaeger indices
curl -X PUT "localhost:9200/_template/jaeger" -H 'Content-Type: application/json' -d'
{
  "index_patterns": ["jaeger-*"],
  "settings": {
    "index.lifecycle.name": "jaeger-policy",
    "index.lifecycle.rollover_alias": "jaeger"
  }
}'

Integration with OpenTelemetry

Key Concepts

OpenTelemetry (OTel) is the recommended instrumentation standard for Jaeger. It provides better support, more features, and vendor-neutral telemetry.

JaegerCollection OptionsApplicationOTel SDKDirect to JaegerOTLPOTel CollectorJaeger CollectorJaeger UIJaegerCollection OptionsApplicationOTel SDKDirect to JaegerOTLPOTel CollectorJaeger CollectorJaeger UI

Python with OpenTelemetry

# Install dependencies
# pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource, SERVICE_NAME

# Configure resource
resource = Resource.create({
    SERVICE_NAME: "my-service",
    "service.version": "1.0.0",
    "deployment.environment": "production"
})

# Set up tracer provider with OTLP exporter to Jaeger
tracer_provider = TracerProvider(resource=resource)
otlp_exporter = OTLPSpanExporter(
    endpoint="http://localhost:4317",  # Jaeger OTLP gRPC endpoint
    insecure=True
)
tracer_provider.add_span_processor(BatchSpanProcessor(otlp_exporter))
trace.set_tracer_provider(tracer_provider)

# Get tracer and create spans
tracer = trace.get_tracer(__name__)

def process_order(order_id):
    with tracer.start_as_current_span("process-order") as span:
        span.set_attribute("order.id", order_id)
        span.set_attribute("order.type", "standard")

        # Add event
        span.add_event("Order validation started")

        validate_order(order_id)

        span.add_event("Order processed successfully")

def validate_order(order_id):
    with tracer.start_as_current_span("validate-order") as span:
        span.set_attribute("order.id", order_id)
        # Validation logic

Go with OpenTelemetry

package main

import (
    "context"
    "log"

    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/attribute"
    "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
    "go.opentelemetry.io/otel/sdk/resource"
    sdktrace "go.opentelemetry.io/otel/sdk/trace"
    semconv "go.opentelemetry.io/otel/semconv/v1.21.0"
)

func initTracer() (*sdktrace.TracerProvider, error) {
    ctx := context.Background()

    // Create OTLP exporter
    exporter, err := otlptracegrpc.New(ctx,
        otlptracegrpc.WithEndpoint("localhost:4317"),
        otlptracegrpc.WithInsecure(),
    )
    if err != nil {
        return nil, err
    }

    // Create resource
    res, err := resource.New(ctx,
        resource.WithAttributes(
            semconv.ServiceName("my-service"),
            semconv.ServiceVersion("1.0.0"),
            attribute.String("environment", "production"),
        ),
    )
    if err != nil {
        return nil, err
    }

    // Create tracer provider
    tp := sdktrace.NewTracerProvider(
        sdktrace.WithBatcher(exporter),
        sdktrace.WithResource(res),
        sdktrace.WithSampler(sdktrace.AlwaysSample()),
    )

    otel.SetTracerProvider(tp)
    return tp, nil
}

func main() {
    tp, err := initTracer()
    if err != nil {
        log.Fatal(err)
    }
    defer tp.Shutdown(context.Background())

    tracer := otel.Tracer("example-tracer")

    ctx, span := tracer.Start(context.Background(), "main-operation")
    defer span.End()

    span.SetAttributes(
        attribute.String("operation.type", "batch"),
        attribute.Int("items.count", 100),
    )

    processItems(ctx)
}

func processItems(ctx context.Context) {
    tracer := otel.Tracer("example-tracer")
    _, span := tracer.Start(ctx, "process-items")
    defer span.End()

    span.AddEvent("Processing started", trace.WithAttributes(
        attribute.Int("batch.size", 100),
    ))
}

Node.js with OpenTelemetry

// tracing.js - Initialisation file
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const { Resource } = require('@opentelemetry/resources');
const { SemanticResourceAttributes } = require('@opentelemetry/semantic-conventions');
const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');

const sdk = new NodeSDK({
    resource: new Resource({
        [SemanticResourceAttributes.SERVICE_NAME]: 'my-service',
        [SemanticResourceAttributes.SERVICE_VERSION]: '1.0.0',
    }),
    traceExporter: new OTLPTraceExporter({
        url: 'http://localhost:4317',
    }),
    instrumentations: [getNodeAutoInstrumentations()],
});

sdk.start();

process.on('SIGTERM', () => {
    sdk.shutdown()
        .then(() => console.log('Tracing terminated'))
        .finally(() => process.exit(0));
});

// app.js - Application code
const { trace } = require('@opentelemetry/api');

const tracer = trace.getTracer('my-service');

async function handleRequest(req) {
    const span = tracer.startSpan('handle-request');

    try {
        span.setAttribute('http.method', req.method);
        span.setAttribute('http.url', req.url);

        const result = await processRequest(req);

        span.setStatus({ code: SpanStatusCode.OK });
        return result;
    } catch (error) {
        span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
        span.recordException(error);
        throw error;
    } finally {
        span.end();
    }
}

OTel Collector with Jaeger Backend

# otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 10s
    send_batch_size: 1000

  memory_limiter:
    check_interval: 1s
    limit_mib: 1000

  attributes:
    actions:
      - key: environment
        value: production
        action: upsert

exporters:
  otlp/jaeger:
    endpoint: jaeger-collector:4317
    tls:
      insecure: true

  # The `logging` exporter was removed in Collector v0.111.0; use `debug`.
  debug:
    verbosity: detailed

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch, attributes]
      exporters: [otlp/jaeger, debug]
# Docker Compose with OTel Collector and Jaeger
version: '3.8'

services:
  otel-collector:
    image: otel/opentelemetry-collector-contrib:latest
    volumes:
      - ./otel-collector-config.yaml:/etc/otelcol/config.yaml
    ports:
      - "4317:4317"   # OTLP gRPC
      - "4318:4318"   # OTLP HTTP
    command: ["--config=/etc/otelcol/config.yaml"]

  jaeger:
    image: jaegertracing/all-in-one:1.52
    ports:
      - "16686:16686"
      - "14250:14250"
    environment:
      - COLLECTOR_OTLP_ENABLED=true

Quick Reference

Essential Commands

Task Command
Run all-in-one docker run -d -p 16686:16686 -p 6831:6831/udp jaegertracing/all-in-one:1.52
Access UI http://localhost:16686
Health check curl http://localhost:14269/
Get services curl http://localhost:16686/api/services
Get trace curl http://localhost:16686/api/traces/{traceId}

Environment Variables

Variable Description Default
SPAN_STORAGE_TYPE Storage backend memory
ES_SERVER_URLS Elasticsearch URLs -
CASSANDRA_SERVERS Cassandra hosts -
COLLECTOR_QUEUE_SIZE Collector queue size 2000
COLLECTOR_NUM_WORKERS Collector workers 50
QUERY_MAX_CLOCK_SKEW_ADJUSTMENT Max clock skew 0s

Span Tags Reference

Tag Description
span.kind client, server, producer, consumer
http.method HTTP method
http.url Full URL
http.status_code Response status
db.type Database type
db.statement Query statement
error Boolean error flag
sampling.priority Force sampling decision

Docker Quick Start Commands

# Development with memory storage
# COLLECTOR_OTLP_ENABLED=true is required before v1.46 to accept OTLP on 4317/4318
docker run -d --name jaeger \
  -e COLLECTOR_OTLP_ENABLED=true \
  -p 6831:6831/udp -p 16686:16686 -p 4317:4317 \
  jaegertracing/all-in-one:1.52

# Production with Elasticsearch
docker run -d --name jaeger-collector \
  -e SPAN_STORAGE_TYPE=elasticsearch \
  -e ES_SERVER_URLS=http://elasticsearch:9200 \
  -p 14268:14268 -p 14250:14250 -p 4317:4317 \
  jaegertracing/jaeger-collector:1.52

# Query service
docker run -d --name jaeger-query \
  -e SPAN_STORAGE_TYPE=elasticsearch \
  -e ES_SERVER_URLS=http://elasticsearch:9200 \
  -p 16686:16686 \
  jaegertracing/jaeger-query:1.52

Common Issues and Solutions

Issue: Traces Not Appearing in UI

Symptoms: Instrumented application runs but no traces visible in Jaeger UI.

Solutions:

# 1. Verify Jaeger is running
curl http://localhost:14269/
curl http://localhost:16686/api/services

# 2. Check agent connectivity (if using agent)
nc -vz localhost 6831

# 3. Check collector connectivity
curl -X POST http://localhost:14268/api/traces \
  -H "Content-Type: application/x-thrift"

# 4. Enable debug logging in application
export JAEGER_REPORTER_LOG_SPANS=true

# 5. Verify service name is set
# Check that JAEGER_SERVICE_NAME or service name in config is correct
# Python - Add console logging for debugging
from jaeger_client import Config
from jaeger_client.reporter import InMemoryReporter

config = Config(
    config={
        'sampler': {'type': 'const', 'param': 1},
        'logging': True,  # Enable logging
    },
    service_name='my-service',
    validate=True,
)

Issue: Missing Spans in Distributed Trace

Symptoms: Parent spans visible but child spans from other services missing.

Solutions:

# Verify context propagation
# Ensure headers are being injected
headers = {}
tracer.inject(span.context, Format.HTTP_HEADERS, headers)
print(f"Injected headers: {headers}")  # Should contain uber-trace-id

# Ensure headers are extracted
span_context = tracer.extract(Format.HTTP_HEADERS, request.headers)
print(f"Extracted context: {span_context}")  # Should not be None
# Check all services use same trace ID format
# Look for uber-trace-id header:
# Format: {trace-id}:{span-id}:{parent-span-id}:{flags}

Issue: High Memory Usage

Symptoms: Jaeger collector or query service consuming excessive memory.

Solutions:

# Reduce collector queue size
collector:
  queue-size: 1000  # Default is 2000
  num-workers: 25   # Default is 50

# Configure batching
reporter:
  batch-size: 100   # Default varies by client
  buffer-flush-interval: 1s
# For Elasticsearch, check index settings
curl http://localhost:9200/_cat/indices/jaeger*?v

# Reduce shard count
ES_NUM_SHARDS=3
ES_NUM_REPLICAS=0  # For development

Issue: Clock Skew Warnings

Symptoms: "Clock skew adjustment" warnings or incorrect span ordering.

Solutions:

# 1. Synchronise clocks across services
# Ensure NTP is running on all hosts

# 2. Increase clock skew tolerance
docker run -d \
  -e QUERY_MAX_CLOCK_SKEW_ADJUSTMENT=5s \
  jaegertracing/jaeger-query:1.52

# 3. Check system time
date
timedatectl status

Issue: Sampling Not Working

Symptoms: All traces sampled or no traces sampled regardless of configuration.

Solutions:

# Verify sampler configuration
config = Config(
    config={
        'sampler': {
            'type': 'probabilistic',
            'param': 0.1,  # 10% sampling
        },
    },
    service_name='my-service',
)

# Check for sampling priority override
# Setting sampling.priority=1 forces sampling
span.set_tag('sampling.priority', 1)  # Forces this trace to be sampled
# For remote sampling, check agent endpoint
curl http://localhost:5778/sampling?service=my-service

Issue: Elasticsearch Connection Errors

Symptoms: Collector fails to connect to Elasticsearch.

Solutions:

# 1. Verify Elasticsearch is accessible
curl http://elasticsearch:9200/_cluster/health

# 2. Check credentials
ES_USERNAME=elastic
ES_PASSWORD=changeme

# 3. Verify index exists
curl http://elasticsearch:9200/_cat/indices/jaeger*

# 4. Create index template if missing
curl -X PUT "http://elasticsearch:9200/_template/jaeger" \
  -H "Content-Type: application/json" \
  -d @jaeger-index-template.json

Issue: Traces Lost During High Load

Symptoms: Traces intermittently missing during traffic spikes.

Solutions:

# Increase collector capacity
collector:
  queue-size: 5000
  num-workers: 100

# Use Kafka buffer for burst handling
environment:
  - SPAN_STORAGE_TYPE=kafka
  - KAFKA_PRODUCER_BROKERS=kafka:9092
# Client-side: increase buffer and reduce sampling
config = Config(
    config={
        'sampler': {'type': 'probabilistic', 'param': 0.01},  # 1%
        'reporter_queue_size': 1000,
        'reporter_batch_size': 200,
    },
    service_name='my-service',
)

Issue: gRPC Connection Failures

Symptoms: "connection refused" or "deadline exceeded" errors with gRPC collector.

Solutions:

# 1. Verify port is open
nc -vz jaeger-collector 14250

# 2. Check TLS configuration
# If using TLS, ensure certificates are valid
openssl s_client -connect jaeger-collector:14250

# 3. Increase timeout
config = Config(
    config={
        'local_agent': {
            'reporting_host': 'localhost',
            'reporting_port': '6831',
        },
        'reporter': {
            'collector_endpoint': 'http://jaeger-collector:14268/api/traces',
        },
    },
)

Related Topics

The following topics complement Jaeger knowledge:

  1. OpenTelemetry - The recommended instrumentation framework for Jaeger, providing vendor-neutral telemetry APIs and SDKs
  2. Prometheus - Metrics collection system often used alongside Jaeger for complete observability
  3. Grafana - Visualisation platform that can display Jaeger traces alongside metrics dashboards
  4. Kubernetes Observability - Deploying and configuring Jaeger in Kubernetes with operators and sidecars
  5. Elasticsearch - Primary production storage backend for Jaeger traces
  6. Service Mesh (Istio) - Automatic tracing instrumentation for microservices without code changes