Jaeger
An open-source distributed tracing system for monitoring and troubleshooting microservices-based architectures.
Jaeger Cheatsheet
An open-source distributed tracing system for monitoring and troubleshooting microservices-based architectures.
Version note (important): This cheatsheet documents Jaeger v1, which reached end-of-life on 31 December 2025. The current release is Jaeger v2 (built on the OpenTelemetry Collector), which ships a single
jaegertracing/jaegerbinary replacing the separate v1 all-in-one/collector/query/ingester images and the Jaeger Agent. Thejaeger-clientinstrumentation libraries shown below were retired in 2022 in favour of the OpenTelemetry SDKs — use OpenTelemetry for all new instrumentation (see the Integration with OpenTelemetry section). The v1 material is retained for existing deployments; migrate to v2 + OTel for anything new.
Overview
Jaeger is a distributed tracing platform originally developed by Uber and now a CNCF graduated project. It helps developers monitor and troubleshoot transactions in complex distributed systems by visualising request flows across services, identifying performance bottlenecks, and analysing root causes of errors.
graph TB
subgraph "Application Layer"
APP1[Service A]
APP2[Service B]
APP3[Service C]
end
subgraph "Jaeger Architecture"
AGENT[Jaeger Agent]
COLLECTOR[Jaeger Collector]
QUERY[Jaeger Query]
UI[Jaeger UI]
end
subgraph "Storage Layer"
STORAGE[(Storage Backend)]
end
APP1 --> AGENT
APP2 --> AGENT
APP3 --> AGENT
AGENT --> COLLECTOR
COLLECTOR --> STORAGE
QUERY --> STORAGE
UI --> QUERY
Core Features
| Feature | Description |
|---|---|
| Distributed Context Propagation | Tracks requests across service boundaries |
| Distributed Transaction Monitoring | Monitors end-to-end request flows |
| Root Cause Analysis | Identifies performance bottlenecks and errors |
| Service Dependency Analysis | Visualises service relationships |
| Performance Optimisation | Measures latency and throughput |
Architecture
Jaeger's architecture consists of several components that work together to collect, store, and visualise traces.
flowchart LR
subgraph "Data Collection"
CLIENT[Instrumented App]
AGENT[Agent<br/>UDP 6831/6832]
end
subgraph "Data Processing"
COLLECTOR[Collector<br/>gRPC 14250<br/>HTTP 14268]
KAFKA[(Kafka<br/>Optional)]
INGESTER[Ingester]
end
subgraph "Storage"
ES[(Elasticsearch)]
CASS[(Cassandra)]
MEM[(Memory)]
end
subgraph "Presentation"
QUERY[Query Service<br/>HTTP 16686]
UI[Web UI]
end
CLIENT -->|Thrift/UDP| AGENT
AGENT -->|gRPC| COLLECTOR
COLLECTOR --> KAFKA
COLLECTOR --> ES
COLLECTOR --> CASS
COLLECTOR --> MEM
KAFKA --> INGESTER
INGESTER --> ES
INGESTER --> CASS
QUERY --> ES
QUERY --> CASS
QUERY --> MEM
UI --> QUERY
Jaeger Agent
The agent is a daemon that listens for spans sent over UDP, batches them, and forwards to the collector. It's typically deployed as a sidecar or DaemonSet.
# Kubernetes DaemonSet deployment
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: jaeger-agent
spec:
selector:
matchLabels:
app: jaeger-agent
template:
metadata:
labels:
app: jaeger-agent
spec:
containers:
- name: jaeger-agent
image: jaegertracing/jaeger-agent:1.52
ports:
- containerPort: 6831
protocol: UDP
name: compact-thrift
- containerPort: 6832
protocol: UDP
name: binary-thrift
- containerPort: 5778
protocol: TCP
name: config
args:
- --reporter.grpc.host-port=jaeger-collector:14250
- --log-level=info
Jaeger Collector
The collector receives traces from agents, validates them, and stores them in a backend storage system.
# Docker Compose collector configuration
jaeger-collector:
image: jaegertracing/jaeger-collector:1.52
ports:
- "14269:14269" # Admin health check
- "14268:14268" # HTTP collector
- "14250:14250" # gRPC collector
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
environment:
- SPAN_STORAGE_TYPE=elasticsearch
- ES_SERVER_URLS=http://elasticsearch:9200
- ES_TAGS_AS_FIELDS_ALL=true
command:
- --collector.queue-size=2000
- --collector.num-workers=50
Jaeger Query Service
The query service exposes APIs to retrieve traces and hosts the web UI.
# Docker Compose query service
jaeger-query:
image: jaegertracing/jaeger-query:1.52
ports:
- "16686:16686" # Web UI
- "16687:16687" # Admin health check
environment:
- SPAN_STORAGE_TYPE=elasticsearch
- ES_SERVER_URLS=http://elasticsearch:9200
depends_on:
- elasticsearch
All-in-One Deployment
For development and testing, use the all-in-one image:
# Run Jaeger all-in-one with in-memory storage
# Note: on Jaeger v1 before 1.46, OTLP ingestion (ports 4317/4318) is
# OFF by default and must be enabled with COLLECTOR_OTLP_ENABLED=true,
# otherwise those ports silently accept nothing. From 1.46 it defaults on.
# The env var is harmless on newer versions so including it aids portability.
docker run -d --name jaeger \
-e COLLECTOR_OTLP_ENABLED=true \
-p 6831:6831/udp \
-p 6832:6832/udp \
-p 5778:5778 \
-p 16686:16686 \
-p 4317:4317 \
-p 4318:4318 \
-p 14250:14250 \
-p 14268:14268 \
-p 14269:14269 \
jaegertracing/all-in-one:1.52
# Access UI at http://localhost:16686
Port Reference
| Port | Protocol | Component | Purpose |
|---|---|---|---|
| 6831 | UDP | Agent | Thrift compact protocol |
| 6832 | UDP | Agent | Thrift binary protocol |
| 5778 | HTTP | Agent | Configuration/sampling |
| 14250 | gRPC | Collector | Model proto |
| 14268 | HTTP | Collector | Jaeger thrift |
| 14269 | HTTP | Collector | Admin/health check |
| 4317 | gRPC | Collector | OTLP gRPC |
| 4318 | HTTP | Collector | OTLP HTTP |
| 16686 | HTTP | Query | Web UI |
| 16687 | HTTP | Query | Admin/health check |
Instrumentation Setup
Key Concepts
Span: A logical unit of work with operation name, start time, duration, and tags.
Trace: A collection of spans forming a directed acyclic graph (DAG).
SpanContext: Carries trace identity across process boundaries.
Tags: Key-value pairs providing metadata about a span.
Logs: Timestamped events within a span.
Python Instrumentation
# Install dependencies
# pip install jaeger-client opentracing
from jaeger_client import Config
from opentracing.ext import tags
from opentracing.propagation import Format
# Initialise tracer
def init_tracer(service_name):
config = Config(
config={
'sampler': {
'type': 'const',
'param': 1, # Sample all traces
},
'local_agent': {
'reporting_host': 'localhost',
'reporting_port': '6831',
},
'logging': True,
},
service_name=service_name,
validate=True,
)
return config.initialize_tracer()
tracer = init_tracer('my-service')
# Create spans
def process_request(request):
with tracer.start_active_span('process-request') as scope:
span = scope.span
span.set_tag(tags.HTTP_METHOD, 'GET')
span.set_tag(tags.HTTP_URL, request.url)
span.set_tag('user.id', request.user_id)
# Log events within span
span.log_kv({'event': 'processing started'})
result = do_work()
span.log_kv({'event': 'processing completed', 'result': result})
span.set_tag(tags.HTTP_STATUS_CODE, 200)
return result
# Propagate context to downstream services
def call_downstream_service(url):
with tracer.start_active_span('call-downstream') as scope:
span = scope.span
# Inject trace context into headers
headers = {}
tracer.inject(span.context, Format.HTTP_HEADERS, headers)
response = requests.get(url, headers=headers)
span.set_tag(tags.HTTP_STATUS_CODE, response.status_code)
return response
# Extract context from incoming request
def handle_incoming_request(request):
# Extract parent span context
span_context = tracer.extract(Format.HTTP_HEADERS, request.headers)
with tracer.start_active_span('handle-request', child_of=span_context) as scope:
# Process request
return process(request)
Go Instrumentation
package main
import (
"context"
"io"
"net/http"
"github.com/opentracing/opentracing-go"
"github.com/opentracing/opentracing-go/ext"
"github.com/uber/jaeger-client-go"
"github.com/uber/jaeger-client-go/config"
)
// Initialise tracer
func initTracer(serviceName string) (opentracing.Tracer, io.Closer, error) {
cfg := config.Configuration{
ServiceName: serviceName,
Sampler: &config.SamplerConfig{
Type: jaeger.SamplerTypeConst,
Param: 1,
},
Reporter: &config.ReporterConfig{
LocalAgentHostPort: "localhost:6831",
LogSpans: true,
},
}
tracer, closer, err := cfg.NewTracer()
if err != nil {
return nil, nil, err
}
opentracing.SetGlobalTracer(tracer)
return tracer, closer, nil
}
func main() {
tracer, closer, err := initTracer("my-service")
if err != nil {
panic(err)
}
defer closer.Close()
// Create root span
span := tracer.StartSpan("main-operation")
defer span.Finish()
ctx := opentracing.ContextWithSpan(context.Background(), span)
processData(ctx)
}
func processData(ctx context.Context) {
// Create child span
span, ctx := opentracing.StartSpanFromContext(ctx, "process-data")
defer span.Finish()
span.SetTag("data.type", "user")
span.LogKV("event", "processing started", "items", 100)
// HTTP client with tracing
callService(ctx, "http://service-b/api/data")
}
func callService(ctx context.Context, url string) {
span, _ := opentracing.StartSpanFromContext(ctx, "http-call")
defer span.Finish()
ext.SpanKindRPCClient.Set(span)
ext.HTTPUrl.Set(span, url)
ext.HTTPMethod.Set(span, "GET")
req, _ := http.NewRequest("GET", url, nil)
// Inject trace context
opentracing.GlobalTracer().Inject(
span.Context(),
opentracing.HTTPHeaders,
opentracing.HTTPHeadersCarrier(req.Header),
)
client := &http.Client{}
resp, err := client.Do(req)
if err != nil {
ext.LogError(span, err)
return
}
defer resp.Body.Close()
ext.HTTPStatusCode.Set(span, uint16(resp.StatusCode))
}
Java Instrumentation
import io.jaegertracing.Configuration;
import io.jaegertracing.Configuration.SamplerConfiguration;
import io.jaegertracing.Configuration.ReporterConfiguration;
import io.opentracing.Scope;
import io.opentracing.Span;
import io.opentracing.Tracer;
import io.opentracing.propagation.Format;
import io.opentracing.propagation.TextMapAdapter;
import io.opentracing.tag.Tags;
public class TracingExample {
private static Tracer initTracer(String serviceName) {
SamplerConfiguration sampler = SamplerConfiguration.fromEnv()
.withType("const")
.withParam(1);
ReporterConfiguration reporter = ReporterConfiguration.fromEnv()
.withLogSpans(true);
return new Configuration(serviceName)
.withSampler(sampler)
.withReporter(reporter)
.getTracer();
}
public static void main(String[] args) {
Tracer tracer = initTracer("my-service");
// Create span with try-with-resources
try (Scope scope = tracer.buildSpan("main-operation")
.withTag("environment", "production")
.startActive(true)) {
Span span = scope.span();
span.log("Processing started");
processRequest(tracer);
span.log("Processing completed");
}
tracer.close();
}
private static void processRequest(Tracer tracer) {
Span parentSpan = tracer.activeSpan();
try (Scope scope = tracer.buildSpan("process-request")
.asChildOf(parentSpan)
.withTag(Tags.SPAN_KIND.getKey(), Tags.SPAN_KIND_SERVER)
.startActive(true)) {
Span span = scope.span();
span.setTag("request.type", "api");
// Inject context for downstream call
Map<String, String> headers = new HashMap<>();
tracer.inject(
span.context(),
Format.Builtin.HTTP_HEADERS,
new TextMapAdapter(headers)
);
// Make HTTP call with headers
callDownstream(headers);
}
}
}
Node.js Instrumentation
const { initTracer } = require('jaeger-client');
const opentracing = require('opentracing');
// Initialise tracer
const config = {
serviceName: 'my-service',
sampler: {
type: 'const',
param: 1,
},
reporter: {
agentHost: 'localhost',
agentPort: 6831,
logSpans: true,
},
};
const options = {
tags: {
'my-service.version': '1.0.0',
},
};
const tracer = initTracer(config, options);
// Create spans
async function processRequest(req, res) {
// Extract context from incoming request
const parentSpanContext = tracer.extract(
opentracing.FORMAT_HTTP_HEADERS,
req.headers
);
const span = tracer.startSpan('process-request', {
childOf: parentSpanContext,
tags: {
[opentracing.Tags.SPAN_KIND]: opentracing.Tags.SPAN_KIND_RPC_SERVER,
[opentracing.Tags.HTTP_METHOD]: req.method,
[opentracing.Tags.HTTP_URL]: req.url,
},
});
try {
span.log({ event: 'processing started' });
const result = await doWork();
span.setTag(opentracing.Tags.HTTP_STATUS_CODE, 200);
span.log({ event: 'processing completed', result });
return result;
} catch (error) {
span.setTag(opentracing.Tags.ERROR, true);
span.log({
event: 'error',
'error.object': error,
message: error.message,
stack: error.stack,
});
throw error;
} finally {
span.finish();
}
}
// Propagate context to downstream service
async function callDownstream(parentSpan, url) {
const span = tracer.startSpan('call-downstream', {
childOf: parentSpan,
tags: {
[opentracing.Tags.SPAN_KIND]: opentracing.Tags.SPAN_KIND_RPC_CLIENT,
[opentracing.Tags.HTTP_URL]: url,
},
});
const headers = {};
tracer.inject(span.context(), opentracing.FORMAT_HTTP_HEADERS, headers);
try {
const response = await fetch(url, { headers });
span.setTag(opentracing.Tags.HTTP_STATUS_CODE, response.status);
return response;
} finally {
span.finish();
}
}
Trace Viewing and Analysis
Using the Jaeger UI
The Jaeger UI provides powerful visualisation and analysis capabilities:
flowchart TB
subgraph "Search Panel"
SERVICE[Select Service]
OPERATION[Select Operation]
TAGS[Filter by Tags]
TIME[Time Range]
LIMIT[Result Limit]
end
subgraph "Results View"
LIST[Trace List]
TIMELINE[Timeline View]
GRAPH[Service Graph]
end
subgraph "Trace Detail"
WATERFALL[Waterfall View]
SPANS[Span Details]
LOGS[Span Logs]
COMPARE[Compare Traces]
end
SERVICE --> LIST
OPERATION --> LIST
TAGS --> LIST
TIME --> LIST
LIMIT --> LIST
LIST --> WATERFALL
LIST --> TIMELINE
LIST --> GRAPH
WATERFALL --> SPANS
WATERFALL --> LOGS
SPANS --> COMPARE
Searching Traces
# Search by service name
# UI: Select service from dropdown
# Search by operation name
# UI: Select operation after choosing service
# Search by tags (key=value)
http.status_code=500
user.id=12345
error=true
# Search by duration
# minDuration: 100ms
# maxDuration: 5s
# Search by trace ID
# Direct lookup: http://localhost:16686/trace/<trace-id>
Common Tag Queries
| Query | Purpose |
|---|---|
error=true |
Find all error traces |
http.status_code=500 |
Find server errors |
http.status_code=404 |
Find not found errors |
http.method=POST |
Find POST requests |
db.type=postgresql |
Find database operations |
kafka.topic=orders |
Find Kafka messages |
Trace Comparison
Compare two traces to identify differences:
- Select first trace and click "Compare"
- Select second trace
- View side-by-side comparison
- Identify timing differences and missing spans
API Queries
# Get all services
curl "http://localhost:16686/api/services"
# Get operations for a service
curl "http://localhost:16686/api/services/my-service/operations"
# Search traces
curl "http://localhost:16686/api/traces?service=my-service&limit=20"
# Get trace by ID
curl "http://localhost:16686/api/traces/<trace-id>"
# Search with tags
curl "http://localhost:16686/api/traces?service=my-service&tags=%7B%22error%22%3A%22true%22%7D"
# Get dependencies
curl "http://localhost:16686/api/dependencies?endTs=$(date +%s)000&lookback=3600000"
Analysing Performance
# Python script to analyse trace data
import requests
import statistics
def analyse_traces(service_name, operation, limit=100):
"""Analyse trace latencies for a service operation."""
url = f"http://localhost:16686/api/traces"
params = {
"service": service_name,
"operation": operation,
"limit": limit
}
response = requests.get(url, params=params)
traces = response.json()['data']
durations = []
for trace in traces:
# Get root span duration
for span in trace['spans']:
if span['operationName'] == operation:
duration_us = span['duration']
durations.append(duration_us / 1000) # Convert to ms
break
if durations:
print(f"Traces analysed: {len(durations)}")
print(f"Min latency: {min(durations):.2f}ms")
print(f"Max latency: {max(durations):.2f}ms")
print(f"Mean latency: {statistics.mean(durations):.2f}ms")
print(f"Median latency: {statistics.median(durations):.2f}ms")
print(f"P95 latency: {sorted(durations)[int(len(durations)*0.95)]:.2f}ms")
print(f"P99 latency: {sorted(durations)[int(len(durations)*0.99)]:.2f}ms")
analyse_traces("api-gateway", "GET /api/orders")
Sampling Strategies
Key Concepts
Sampling determines which traces to record. Different strategies balance visibility with overhead:
| Strategy | Description | Use Case |
|---|---|---|
| Const | Sample all (1) or none (0) | Development/debugging |
| Probabilistic | Sample percentage of traces | Production baseline |
| Rate Limiting | Fixed traces per second | Cost control |
| Remote | Server-controlled sampling | Dynamic adjustment |
| Adaptive | Automatic adjustment | High-traffic systems |
Configuration Examples
# Const sampler - sample everything
sampler:
type: const
param: 1 # 1 = sample all, 0 = sample none
# Probabilistic sampler - 10% of traces
sampler:
type: probabilistic
param: 0.1
# Rate limiting sampler - 2 traces per second
sampler:
type: ratelimiting
param: 2
# Remote sampler - fetch from agent
sampler:
type: remote
param: 1
sampling_server_url: http://localhost:5778/sampling
Remote Sampling Configuration
Configure the collector to serve sampling strategies:
// sampling-strategies.json
{
"service_strategies": [
{
"service": "api-gateway",
"type": "probabilistic",
"param": 0.5
},
{
"service": "payment-service",
"type": "probabilistic",
"param": 1.0
},
{
"service": "logging-service",
"type": "ratelimiting",
"param": 10
}
],
"default_strategy": {
"type": "probabilistic",
"param": 0.1
},
"per_operation_strategies": [
{
"service": "api-gateway",
"operation_strategies": [
{
"operation": "GET /health",
"type": "probabilistic",
"param": 0.001
},
{
"operation": "POST /api/orders",
"type": "probabilistic",
"param": 1.0
}
]
}
]
}
# Start collector with sampling strategies
docker run -d \
-p 14268:14268 \
-p 5778:5778 \
-v $(pwd)/sampling-strategies.json:/etc/jaeger/sampling-strategies.json \
jaegertracing/jaeger-collector:1.52 \
--sampling.strategies-file=/etc/jaeger/sampling-strategies.json
Adaptive Sampling
# Collector configuration for adaptive sampling
collector:
sampling:
strategies-file: /etc/jaeger/sampling-strategies.json
strategies-reload-interval: 30s
# Adaptive sampling (Jaeger 1.27+)
sampling:
default_strategy:
type: adaptive
adaptive:
sampling_store: memory
target_samples_per_second: 1.0
delta_tolerance: 0.3
buckets_for_calculation: 1
calculation_interval: 1m
Priority-Based Sampling
// Go - Priority tag for guaranteed sampling
func createHighPrioritySpan(tracer opentracing.Tracer) {
span := tracer.StartSpan("critical-operation")
// Set sampling priority to ensure trace is recorded
ext.SamplingPriority.Set(span, 1)
defer span.Finish()
}
// In error handling
func handleError(span opentracing.Span, err error) {
// Force sampling on errors
ext.SamplingPriority.Set(span, 1)
ext.LogError(span, err)
}
Storage Backends
Supported Storage Options
graph TB
JAEGER[Jaeger Collector]
subgraph "Production Storage"
ES[Elasticsearch]
CASS[Cassandra]
KAFKA[Kafka + Flink]
end
subgraph "Development Storage"
MEM[Memory]
BADGER[Badger]
end
subgraph "Cloud Storage"
ESCLOUD[Elasticsearch Service]
CASSCLOUD[Cassandra as a Service]
end
JAEGER --> ES
JAEGER --> CASS
JAEGER --> KAFKA
JAEGER --> MEM
JAEGER --> BADGER
JAEGER --> ESCLOUD
JAEGER --> CASSCLOUD
Elasticsearch Configuration
# Docker Compose with Elasticsearch
version: '3.8'
services:
elasticsearch:
image: docker.elastic.co/elasticsearch/elasticsearch:8.11.0
environment:
- discovery.type=single-node
- xpack.security.enabled=false
- "ES_JAVA_OPTS=-Xms1g -Xmx1g"
ports:
- "9200:9200"
volumes:
- esdata:/usr/share/elasticsearch/data
jaeger-collector:
image: jaegertracing/jaeger-collector:1.52
environment:
- SPAN_STORAGE_TYPE=elasticsearch
- ES_SERVER_URLS=http://elasticsearch:9200
- ES_NUM_SHARDS=5
- ES_NUM_REPLICAS=1
- ES_INDEX_PREFIX=jaeger
- ES_TAGS_AS_FIELDS_ALL=true
- ES_BULK_SIZE=5000000
- ES_BULK_WORKERS=2
- ES_BULK_ACTIONS=1000
- ES_BULK_FLUSH_INTERVAL=200ms
ports:
- "14268:14268"
- "14250:14250"
jaeger-query:
image: jaegertracing/jaeger-query:1.52
environment:
- SPAN_STORAGE_TYPE=elasticsearch
- ES_SERVER_URLS=http://elasticsearch:9200
- ES_INDEX_PREFIX=jaeger
ports:
- "16686:16686"
volumes:
esdata:
Cassandra Configuration
# Docker Compose with Cassandra
version: '3.8'
services:
cassandra:
image: cassandra:4.1
ports:
- "9042:9042"
volumes:
- cassdata:/var/lib/cassandra
environment:
- CASSANDRA_DC=dc1
- CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch
cassandra-schema:
image: jaegertracing/jaeger-cassandra-schema:1.52
depends_on:
- cassandra
environment:
- CQLSH_HOST=cassandra
- DATACENTER=dc1
- MODE=prod
jaeger-collector:
image: jaegertracing/jaeger-collector:1.52
environment:
- SPAN_STORAGE_TYPE=cassandra
- CASSANDRA_SERVERS=cassandra
- CASSANDRA_KEYSPACE=jaeger_v1_dc1
ports:
- "14268:14268"
- "14250:14250"
depends_on:
- cassandra-schema
jaeger-query:
image: jaegertracing/jaeger-query:1.52
environment:
- SPAN_STORAGE_TYPE=cassandra
- CASSANDRA_SERVERS=cassandra
- CASSANDRA_KEYSPACE=jaeger_v1_dc1
ports:
- "16686:16686"
volumes:
cassdata:
Kafka Streaming Pipeline
# High-throughput setup with Kafka
version: '3.8'
services:
kafka:
image: confluentinc/cp-kafka:7.5.0
environment:
KAFKA_BROKER_ID: 1
KAFKA_ZOOKEEPER_CONNECT: zookeeper:2181
KAFKA_ADVERTISED_LISTENERS: PLAINTEXT://kafka:9092
KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1
jaeger-collector:
image: jaegertracing/jaeger-collector:1.52
environment:
- SPAN_STORAGE_TYPE=kafka
- KAFKA_PRODUCER_BROKERS=kafka:9092
- KAFKA_PRODUCER_TOPIC=jaeger-spans
ports:
- "14268:14268"
jaeger-ingester:
image: jaegertracing/jaeger-ingester:1.52
environment:
- SPAN_STORAGE_TYPE=elasticsearch
- KAFKA_CONSUMER_BROKERS=kafka:9092
- KAFKA_CONSUMER_TOPIC=jaeger-spans
- KAFKA_CONSUMER_GROUP=jaeger-ingester
- ES_SERVER_URLS=http://elasticsearch:9200
- INGESTER_PARALLELISM=1000
- INGESTER_DEADLOCKINTERVAL=5m
Badger (Local Storage)
# For development/testing with persistent storage
docker run -d \
--name jaeger \
-p 16686:16686 \
-p 6831:6831/udp \
-v jaeger_data:/badger \
-e SPAN_STORAGE_TYPE=badger \
-e BADGER_EPHEMERAL=false \
-e BADGER_DIRECTORY_VALUE=/badger/data \
-e BADGER_DIRECTORY_KEY=/badger/key \
jaegertracing/all-in-one:1.52
Index Management
# Elasticsearch index lifecycle management
# Create ILM policy
curl -X PUT "localhost:9200/_ilm/policy/jaeger-policy" -H 'Content-Type: application/json' -d'
{
"policy": {
"phases": {
"hot": {
"min_age": "0ms",
"actions": {
"rollover": {
"max_age": "1d",
"max_size": "50gb"
}
}
},
"delete": {
"min_age": "7d",
"actions": {
"delete": {}
}
}
}
}
}'
# Apply to Jaeger indices
curl -X PUT "localhost:9200/_template/jaeger" -H 'Content-Type: application/json' -d'
{
"index_patterns": ["jaeger-*"],
"settings": {
"index.lifecycle.name": "jaeger-policy",
"index.lifecycle.rollover_alias": "jaeger"
}
}'
Integration with OpenTelemetry
Key Concepts
OpenTelemetry (OTel) is the recommended instrumentation standard for Jaeger. It provides better support, more features, and vendor-neutral telemetry.
flowchart LR
subgraph "Application"
OTEL_SDK[OTel SDK]
end
subgraph "Collection Options"
DIRECT[Direct to Jaeger<br/>OTLP]
COLLECTOR[OTel Collector]
end
subgraph "Jaeger"
JAEGER_COL[Jaeger Collector]
JAEGER_UI[Jaeger UI]
end
OTEL_SDK --> DIRECT
OTEL_SDK --> COLLECTOR
DIRECT --> JAEGER_COL
COLLECTOR --> JAEGER_COL
JAEGER_COL --> JAEGER_UI
Python with OpenTelemetry
# Install dependencies
# pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource, SERVICE_NAME
# Configure resource
resource = Resource.create({
SERVICE_NAME: "my-service",
"service.version": "1.0.0",
"deployment.environment": "production"
})
# Set up tracer provider with OTLP exporter to Jaeger
tracer_provider = TracerProvider(resource=resource)
otlp_exporter = OTLPSpanExporter(
endpoint="http://localhost:4317", # Jaeger OTLP gRPC endpoint
insecure=True
)
tracer_provider.add_span_processor(BatchSpanProcessor(otlp_exporter))
trace.set_tracer_provider(tracer_provider)
# Get tracer and create spans
tracer = trace.get_tracer(__name__)
def process_order(order_id):
with tracer.start_as_current_span("process-order") as span:
span.set_attribute("order.id", order_id)
span.set_attribute("order.type", "standard")
# Add event
span.add_event("Order validation started")
validate_order(order_id)
span.add_event("Order processed successfully")
def validate_order(order_id):
with tracer.start_as_current_span("validate-order") as span:
span.set_attribute("order.id", order_id)
# Validation logic
Go with OpenTelemetry
package main
import (
"context"
"log"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
semconv "go.opentelemetry.io/otel/semconv/v1.21.0"
)
func initTracer() (*sdktrace.TracerProvider, error) {
ctx := context.Background()
// Create OTLP exporter
exporter, err := otlptracegrpc.New(ctx,
otlptracegrpc.WithEndpoint("localhost:4317"),
otlptracegrpc.WithInsecure(),
)
if err != nil {
return nil, err
}
// Create resource
res, err := resource.New(ctx,
resource.WithAttributes(
semconv.ServiceName("my-service"),
semconv.ServiceVersion("1.0.0"),
attribute.String("environment", "production"),
),
)
if err != nil {
return nil, err
}
// Create tracer provider
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exporter),
sdktrace.WithResource(res),
sdktrace.WithSampler(sdktrace.AlwaysSample()),
)
otel.SetTracerProvider(tp)
return tp, nil
}
func main() {
tp, err := initTracer()
if err != nil {
log.Fatal(err)
}
defer tp.Shutdown(context.Background())
tracer := otel.Tracer("example-tracer")
ctx, span := tracer.Start(context.Background(), "main-operation")
defer span.End()
span.SetAttributes(
attribute.String("operation.type", "batch"),
attribute.Int("items.count", 100),
)
processItems(ctx)
}
func processItems(ctx context.Context) {
tracer := otel.Tracer("example-tracer")
_, span := tracer.Start(ctx, "process-items")
defer span.End()
span.AddEvent("Processing started", trace.WithAttributes(
attribute.Int("batch.size", 100),
))
}
Node.js with OpenTelemetry
// tracing.js - Initialisation file
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const { Resource } = require('@opentelemetry/resources');
const { SemanticResourceAttributes } = require('@opentelemetry/semantic-conventions');
const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: 'my-service',
[SemanticResourceAttributes.SERVICE_VERSION]: '1.0.0',
}),
traceExporter: new OTLPTraceExporter({
url: 'http://localhost:4317',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
process.on('SIGTERM', () => {
sdk.shutdown()
.then(() => console.log('Tracing terminated'))
.finally(() => process.exit(0));
});
// app.js - Application code
const { trace } = require('@opentelemetry/api');
const tracer = trace.getTracer('my-service');
async function handleRequest(req) {
const span = tracer.startSpan('handle-request');
try {
span.setAttribute('http.method', req.method);
span.setAttribute('http.url', req.url);
const result = await processRequest(req);
span.setStatus({ code: SpanStatusCode.OK });
return result;
} catch (error) {
span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
span.recordException(error);
throw error;
} finally {
span.end();
}
}
OTel Collector with Jaeger Backend
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 10s
send_batch_size: 1000
memory_limiter:
check_interval: 1s
limit_mib: 1000
attributes:
actions:
- key: environment
value: production
action: upsert
exporters:
otlp/jaeger:
endpoint: jaeger-collector:4317
tls:
insecure: true
# The `logging` exporter was removed in Collector v0.111.0; use `debug`.
debug:
verbosity: detailed
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch, attributes]
exporters: [otlp/jaeger, debug]
# Docker Compose with OTel Collector and Jaeger
version: '3.8'
services:
otel-collector:
image: otel/opentelemetry-collector-contrib:latest
volumes:
- ./otel-collector-config.yaml:/etc/otelcol/config.yaml
ports:
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
command: ["--config=/etc/otelcol/config.yaml"]
jaeger:
image: jaegertracing/all-in-one:1.52
ports:
- "16686:16686"
- "14250:14250"
environment:
- COLLECTOR_OTLP_ENABLED=true
Quick Reference
Essential Commands
| Task | Command |
|---|---|
| Run all-in-one | docker run -d -p 16686:16686 -p 6831:6831/udp jaegertracing/all-in-one:1.52 |
| Access UI | http://localhost:16686 |
| Health check | curl http://localhost:14269/ |
| Get services | curl http://localhost:16686/api/services |
| Get trace | curl http://localhost:16686/api/traces/{traceId} |
Environment Variables
| Variable | Description | Default |
|---|---|---|
SPAN_STORAGE_TYPE |
Storage backend | memory |
ES_SERVER_URLS |
Elasticsearch URLs | - |
CASSANDRA_SERVERS |
Cassandra hosts | - |
COLLECTOR_QUEUE_SIZE |
Collector queue size | 2000 |
COLLECTOR_NUM_WORKERS |
Collector workers | 50 |
QUERY_MAX_CLOCK_SKEW_ADJUSTMENT |
Max clock skew | 0s |
Span Tags Reference
| Tag | Description |
|---|---|
span.kind |
client, server, producer, consumer |
http.method |
HTTP method |
http.url |
Full URL |
http.status_code |
Response status |
db.type |
Database type |
db.statement |
Query statement |
error |
Boolean error flag |
sampling.priority |
Force sampling decision |
Docker Quick Start Commands
# Development with memory storage
# COLLECTOR_OTLP_ENABLED=true is required before v1.46 to accept OTLP on 4317/4318
docker run -d --name jaeger \
-e COLLECTOR_OTLP_ENABLED=true \
-p 6831:6831/udp -p 16686:16686 -p 4317:4317 \
jaegertracing/all-in-one:1.52
# Production with Elasticsearch
docker run -d --name jaeger-collector \
-e SPAN_STORAGE_TYPE=elasticsearch \
-e ES_SERVER_URLS=http://elasticsearch:9200 \
-p 14268:14268 -p 14250:14250 -p 4317:4317 \
jaegertracing/jaeger-collector:1.52
# Query service
docker run -d --name jaeger-query \
-e SPAN_STORAGE_TYPE=elasticsearch \
-e ES_SERVER_URLS=http://elasticsearch:9200 \
-p 16686:16686 \
jaegertracing/jaeger-query:1.52
Common Issues and Solutions
Issue: Traces Not Appearing in UI
Symptoms: Instrumented application runs but no traces visible in Jaeger UI.
Solutions:
# 1. Verify Jaeger is running
curl http://localhost:14269/
curl http://localhost:16686/api/services
# 2. Check agent connectivity (if using agent)
nc -vz localhost 6831
# 3. Check collector connectivity
curl -X POST http://localhost:14268/api/traces \
-H "Content-Type: application/x-thrift"
# 4. Enable debug logging in application
export JAEGER_REPORTER_LOG_SPANS=true
# 5. Verify service name is set
# Check that JAEGER_SERVICE_NAME or service name in config is correct
# Python - Add console logging for debugging
from jaeger_client import Config
from jaeger_client.reporter import InMemoryReporter
config = Config(
config={
'sampler': {'type': 'const', 'param': 1},
'logging': True, # Enable logging
},
service_name='my-service',
validate=True,
)
Issue: Missing Spans in Distributed Trace
Symptoms: Parent spans visible but child spans from other services missing.
Solutions:
# Verify context propagation
# Ensure headers are being injected
headers = {}
tracer.inject(span.context, Format.HTTP_HEADERS, headers)
print(f"Injected headers: {headers}") # Should contain uber-trace-id
# Ensure headers are extracted
span_context = tracer.extract(Format.HTTP_HEADERS, request.headers)
print(f"Extracted context: {span_context}") # Should not be None
# Check all services use same trace ID format
# Look for uber-trace-id header:
# Format: {trace-id}:{span-id}:{parent-span-id}:{flags}
Issue: High Memory Usage
Symptoms: Jaeger collector or query service consuming excessive memory.
Solutions:
# Reduce collector queue size
collector:
queue-size: 1000 # Default is 2000
num-workers: 25 # Default is 50
# Configure batching
reporter:
batch-size: 100 # Default varies by client
buffer-flush-interval: 1s
# For Elasticsearch, check index settings
curl http://localhost:9200/_cat/indices/jaeger*?v
# Reduce shard count
ES_NUM_SHARDS=3
ES_NUM_REPLICAS=0 # For development
Issue: Clock Skew Warnings
Symptoms: "Clock skew adjustment" warnings or incorrect span ordering.
Solutions:
# 1. Synchronise clocks across services
# Ensure NTP is running on all hosts
# 2. Increase clock skew tolerance
docker run -d \
-e QUERY_MAX_CLOCK_SKEW_ADJUSTMENT=5s \
jaegertracing/jaeger-query:1.52
# 3. Check system time
date
timedatectl status
Issue: Sampling Not Working
Symptoms: All traces sampled or no traces sampled regardless of configuration.
Solutions:
# Verify sampler configuration
config = Config(
config={
'sampler': {
'type': 'probabilistic',
'param': 0.1, # 10% sampling
},
},
service_name='my-service',
)
# Check for sampling priority override
# Setting sampling.priority=1 forces sampling
span.set_tag('sampling.priority', 1) # Forces this trace to be sampled
# For remote sampling, check agent endpoint
curl http://localhost:5778/sampling?service=my-service
Issue: Elasticsearch Connection Errors
Symptoms: Collector fails to connect to Elasticsearch.
Solutions:
# 1. Verify Elasticsearch is accessible
curl http://elasticsearch:9200/_cluster/health
# 2. Check credentials
ES_USERNAME=elastic
ES_PASSWORD=changeme
# 3. Verify index exists
curl http://elasticsearch:9200/_cat/indices/jaeger*
# 4. Create index template if missing
curl -X PUT "http://elasticsearch:9200/_template/jaeger" \
-H "Content-Type: application/json" \
-d @jaeger-index-template.json
Issue: Traces Lost During High Load
Symptoms: Traces intermittently missing during traffic spikes.
Solutions:
# Increase collector capacity
collector:
queue-size: 5000
num-workers: 100
# Use Kafka buffer for burst handling
environment:
- SPAN_STORAGE_TYPE=kafka
- KAFKA_PRODUCER_BROKERS=kafka:9092
# Client-side: increase buffer and reduce sampling
config = Config(
config={
'sampler': {'type': 'probabilistic', 'param': 0.01}, # 1%
'reporter_queue_size': 1000,
'reporter_batch_size': 200,
},
service_name='my-service',
)
Issue: gRPC Connection Failures
Symptoms: "connection refused" or "deadline exceeded" errors with gRPC collector.
Solutions:
# 1. Verify port is open
nc -vz jaeger-collector 14250
# 2. Check TLS configuration
# If using TLS, ensure certificates are valid
openssl s_client -connect jaeger-collector:14250
# 3. Increase timeout
config = Config(
config={
'local_agent': {
'reporting_host': 'localhost',
'reporting_port': '6831',
},
'reporter': {
'collector_endpoint': 'http://jaeger-collector:14268/api/traces',
},
},
)
Related Topics
The following topics complement Jaeger knowledge:
- OpenTelemetry - The recommended instrumentation framework for Jaeger, providing vendor-neutral telemetry APIs and SDKs
- Prometheus - Metrics collection system often used alongside Jaeger for complete observability
- Grafana - Visualisation platform that can display Jaeger traces alongside metrics dashboards
- Kubernetes Observability - Deploying and configuring Jaeger in Kubernetes with operators and sidecars
- Elasticsearch - Primary production storage backend for Jaeger traces
- Service Mesh (Istio) - Automatic tracing instrumentation for microservices without code changes