OpenTelemetry
A vendor-neutral observability framework for generating, collecting, and exporting telemetry data (traces, metrics, and logs) from cloud-native applications.
OpenTelemetry Cheatsheet
A vendor-neutral observability framework for generating, collecting, and exporting telemetry data (traces, metrics, and logs) from cloud-native applications.
Overview
OpenTelemetry (OTel) provides a standardised approach to instrumenting applications and infrastructure for observability. It combines the best of OpenTracing and OpenCensus into a single, unified framework.
graph TB
subgraph "Application Layer"
APP[Application Code]
AUTO[Auto Instrumentation]
MANUAL[Manual Instrumentation]
SDK[OTel SDK]
end
subgraph "Collection Layer"
COLLECTOR[OTel Collector]
RECEIVERS[Receivers]
PROCESSORS[Processors]
EXPORTERS[Exporters]
end
subgraph "Backend Layer"
JAEGER[Jaeger]
PROMETHEUS[Prometheus]
ZIPKIN[Zipkin]
CUSTOM[Custom Backend]
end
APP --> AUTO
APP --> MANUAL
AUTO --> SDK
MANUAL --> SDK
SDK --> COLLECTOR
COLLECTOR --> RECEIVERS
RECEIVERS --> PROCESSORS
PROCESSORS --> EXPORTERS
EXPORTERS --> JAEGER
EXPORTERS --> PROMETHEUS
EXPORTERS --> ZIPKIN
EXPORTERS --> CUSTOM
Core Components
| Component | Description |
|---|---|
| API | Defines how to generate telemetry data |
| SDK | Implements the API with configuration and processing |
| Collector | Receives, processes, and exports telemetry data |
| OTLP | OpenTelemetry Protocol - native wire format |
Traces, Metrics, and Logs
Key Concepts
Traces represent the journey of a request through a distributed system:
- Span: A single unit of work with a name, start time, duration, and attributes
- Trace: A collection of spans forming a directed acyclic graph (DAG)
- SpanContext: Immutable data carried across process boundaries
Metrics capture measurements about a service at runtime:
- Counter: Monotonically increasing value (e.g., requests served)
- Gauge: Point-in-time value (e.g., current temperature)
- Histogram: Distribution of values (e.g., request latencies)
Logs provide timestamped records of discrete events:
- Correlated with traces via trace context
- Include structured attributes and severity levels
Common Patterns
# Python - Creating a trace with spans
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
# Set up the tracer provider
trace.set_tracer_provider(TracerProvider())
tracer = trace.get_tracer(__name__)
# Create spans
with tracer.start_as_current_span("parent-operation") as parent:
parent.set_attribute("user.id", "12345")
with tracer.start_as_current_span("child-operation") as child:
child.add_event("processing started")
# Do work here
child.set_status(trace.StatusCode.OK)
# Python - Recording metrics
from opentelemetry import metrics
from opentelemetry.sdk.metrics import MeterProvider
metrics.set_meter_provider(MeterProvider())
meter = metrics.get_meter(__name__)
# Create instruments
request_counter = meter.create_counter(
name="http_requests_total",
description="Total HTTP requests",
unit="1"
)
request_duration = meter.create_histogram(
name="http_request_duration_seconds",
description="HTTP request duration",
unit="s"
)
# Record measurements
request_counter.add(1, {"method": "GET", "status": "200"})
request_duration.record(0.125, {"method": "GET", "endpoint": "/api/users"})
// JavaScript/Node.js - Creating spans
const { trace } = require('@opentelemetry/api');
const tracer = trace.getTracer('my-service');
async function handleRequest(req, res) {
const span = tracer.startSpan('handle-request');
try {
span.setAttribute('http.method', req.method);
span.setAttribute('http.url', req.url);
const result = await processRequest(req);
span.setStatus({ code: SpanStatusCode.OK });
return result;
} catch (error) {
span.setStatus({
code: SpanStatusCode.ERROR,
message: error.message
});
span.recordException(error);
throw error;
} finally {
span.end();
}
}
Examples
// Go - Complete tracing example
package main
import (
"context"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/trace"
)
func main() {
tracer := otel.Tracer("example-service")
ctx, span := tracer.Start(context.Background(), "main-operation",
trace.WithAttributes(
attribute.String("environment", "production"),
attribute.Int("version", 1),
),
)
defer span.End()
processData(ctx)
}
func processData(ctx context.Context) {
tracer := otel.Tracer("example-service")
_, span := tracer.Start(ctx, "process-data")
defer span.End()
span.AddEvent("data processing complete", trace.WithAttributes(
attribute.Int("items.processed", 100),
))
}
Instrumentation
Key Concepts
Auto-instrumentation: Automatically instruments common libraries and frameworks without code changes.
Manual instrumentation: Explicitly adding spans and metrics to your code for custom business logic.
Semantic conventions: Standardised attribute names ensuring consistency across services.
Auto-Instrumentation Setup
# Python auto-instrumentation
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install
# Run with auto-instrumentation
opentelemetry-instrument \
--traces_exporter otlp \
--metrics_exporter otlp \
--service_name my-service \
python app.py
# Java auto-instrumentation
# Download the agent
curl -L -O https://github.com/open-telemetry/opentelemetry-java-instrumentation/releases/latest/download/opentelemetry-javaagent.jar
# Run with agent
java -javaagent:opentelemetry-javaagent.jar \
-Dotel.service.name=my-service \
-Dotel.exporter.otlp.endpoint=http://collector:4317 \
-jar myapp.jar
# Node.js auto-instrumentation
npm install @opentelemetry/auto-instrumentations-node
# In your application entry point
node --require @opentelemetry/auto-instrumentations-node/register app.js
Manual Instrumentation Examples
# Python - Manual instrumentation with decorators
from opentelemetry import trace
from functools import wraps
tracer = trace.get_tracer(__name__)
def traced(span_name=None):
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
name = span_name or func.__name__
with tracer.start_as_current_span(name) as span:
span.set_attribute("function.name", func.__name__)
return func(*args, **kwargs)
return wrapper
return decorator
@traced("database-query")
def fetch_user(user_id):
# Database query logic
pass
// Java - Manual instrumentation
import io.opentelemetry.api.GlobalOpenTelemetry;
import io.opentelemetry.api.trace.Span;
import io.opentelemetry.api.trace.Tracer;
import io.opentelemetry.context.Scope;
public class UserService {
private static final Tracer tracer =
GlobalOpenTelemetry.getTracer("user-service");
public User getUser(String userId) {
Span span = tracer.spanBuilder("get-user")
.setAttribute("user.id", userId)
.startSpan();
try (Scope scope = span.makeCurrent()) {
User user = userRepository.findById(userId);
span.setAttribute("user.found", user != null);
return user;
} catch (Exception e) {
span.recordException(e);
throw e;
} finally {
span.end();
}
}
}
SDK Configuration
Key Concepts
The SDK provides the implementation of the OpenTelemetry API with:
- Resource: Describes the entity producing telemetry
- Sampler: Determines which traces to record
- SpanProcessor: Processes spans before export
- Exporter: Sends telemetry to backends
Environment Variables
# Core configuration
export OTEL_SERVICE_NAME=my-service
export OTEL_RESOURCE_ATTRIBUTES=service.version=1.0.0,deployment.environment.name=production
# Exporter configuration
export OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
export OTEL_EXPORTER_OTLP_HEADERS=Authorization=Bearer token123
# Trace configuration
export OTEL_TRACES_EXPORTER=otlp
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.1
# Metrics configuration
export OTEL_METRICS_EXPORTER=otlp
export OTEL_METRIC_EXPORT_INTERVAL=60000
# Logs configuration
export OTEL_LOGS_EXPORTER=otlp
# Propagation
export OTEL_PROPAGATORS=tracecontext,baggage
Programmatic Configuration
# Python - Complete SDK setup
from opentelemetry import trace, metrics
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource, SERVICE_NAME
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.exporter.otlp.proto.grpc.metric_exporter import OTLPMetricExporter
# Define resource
resource = Resource.create({
SERVICE_NAME: "my-service",
"service.version": "1.0.0",
"deployment.environment.name": "production"
})
# Configure tracing
tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317"))
)
trace.set_tracer_provider(tracer_provider)
# Configure metrics
metric_reader = PeriodicExportingMetricReader(
OTLPMetricExporter(endpoint="http://collector:4317"),
export_interval_millis=60000
)
meter_provider = MeterProvider(resource=resource, metric_readers=[metric_reader])
metrics.set_meter_provider(meter_provider)
// Node.js - SDK configuration
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { resourceFromAttributes } = require('@opentelemetry/resources');
const { ATTR_SERVICE_NAME, ATTR_SERVICE_VERSION } = require('@opentelemetry/semantic-conventions');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const { OTLPMetricExporter } = require('@opentelemetry/exporter-metrics-otlp-grpc');
const { PeriodicExportingMetricReader } = require('@opentelemetry/sdk-metrics');
const sdk = new NodeSDK({
resource: resourceFromAttributes({
[ATTR_SERVICE_NAME]: 'my-service',
[ATTR_SERVICE_VERSION]: '1.0.0',
}),
traceExporter: new OTLPTraceExporter({
url: 'http://collector:4317',
}),
metricReader: new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({
url: 'http://collector:4317',
}),
exportIntervalMillis: 60000,
}),
});
sdk.start();
process.on('SIGTERM', () => {
sdk.shutdown()
.then(() => console.log('SDK shut down successfully'))
.catch((error) => console.error('Error shutting down SDK', error))
.finally(() => process.exit(0));
});
Sampling Strategies
| Sampler | Description | Use Case |
|---|---|---|
always_on |
Sample all traces | Development/debugging |
always_off |
Sample no traces | Disable tracing |
traceidratio |
Sample percentage of traces | Production cost control |
parentbased_always_on |
Follow parent decision, default on | Distributed systems |
parentbased_traceidratio |
Follow parent, ratio for roots | Production default |
Collector Setup and Pipelines
Key Concepts
The OpenTelemetry Collector is a vendor-agnostic proxy that:
- Receives telemetry from multiple sources
- Processes and transforms data
- Exports to one or more destinations
graph LR
subgraph "Receivers"
OTLP_R[OTLP]
JAEGER_R[Jaeger]
PROM_R[Prometheus]
end
subgraph "Processors"
BATCH[Batch]
FILTER[Filter]
ATTR[Attributes]
TAIL[Tail Sampling]
end
subgraph "Exporters"
OTLP_E[OTLP]
JAEGER_E[Jaeger]
PROM_E[Prometheus]
end
OTLP_R --> BATCH
JAEGER_R --> BATCH
PROM_R --> BATCH
BATCH --> FILTER
FILTER --> ATTR
ATTR --> TAIL
TAIL --> OTLP_E
TAIL --> JAEGER_E
TAIL --> PROM_E
Collector Configuration
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
jaeger:
protocols:
thrift_http:
endpoint: 0.0.0.0:14268
grpc:
endpoint: 0.0.0.0:14250
prometheus:
config:
scrape_configs:
- job_name: 'otel-collector'
scrape_interval: 10s
static_configs:
- targets: ['localhost:8888']
processors:
batch:
timeout: 10s
send_batch_size: 1000
send_batch_max_size: 1500
memory_limiter:
check_interval: 1s
limit_mib: 1000
spike_limit_mib: 200
attributes:
actions:
- key: environment
value: production
action: upsert
- key: sensitive.data
action: delete
filter:
error_mode: ignore
traces:
span:
- 'attributes["http.route"] == "/health"'
- 'name == "internal-healthcheck"'
tail_sampling:
decision_wait: 10s
num_traces: 100000
policies:
- name: errors
type: status_code
status_code:
status_codes:
- ERROR
- name: slow-traces
type: latency
latency:
threshold_ms: 1000
- name: probabilistic
type: probabilistic
probabilistic:
sampling_percentage: 10
exporters:
otlp:
endpoint: tempo:4317
tls:
insecure: true
# Jaeger v2 ingests OTLP natively. The Collector's native `jaeger`
# exporter was removed (last shipped in v0.85.0); export to Jaeger
# over OTLP instead.
otlp/jaeger:
endpoint: jaeger:4317
tls:
insecure: true
prometheus:
endpoint: 0.0.0.0:8889
namespace: otel
const_labels:
service: my-service
# The `logging` exporter was removed in v0.111.0; use `debug`.
debug:
verbosity: detailed
extensions:
health_check:
endpoint: 0.0.0.0:13133
zpages:
endpoint: 0.0.0.0:55679
service:
extensions: [health_check, zpages]
pipelines:
traces:
receivers: [otlp, jaeger]
processors: [memory_limiter, batch, attributes, tail_sampling]
exporters: [otlp, otlp/jaeger]
metrics:
receivers: [otlp, prometheus]
processors: [memory_limiter, batch]
exporters: [prometheus]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp]
Deployment Options
# Docker deployment
docker run -d \
-p 4317:4317 \
-p 4318:4318 \
-p 8889:8889 \
-v $(pwd)/otel-collector-config.yaml:/etc/otelcol/config.yaml \
otel/opentelemetry-collector-contrib:latest
# Kubernetes deployment
kubectl apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
name: otel-collector
spec:
replicas: 2
selector:
matchLabels:
app: otel-collector
template:
metadata:
labels:
app: otel-collector
spec:
containers:
- name: collector
image: otel/opentelemetry-collector-contrib:latest
# The -contrib image defaults to /etc/otelcol-contrib/config.yaml,
# so pass --config explicitly to load the mounted config below.
command: ["/otelcol-contrib"]
args: ["--config=/etc/otelcol/config.yaml"]
ports:
- containerPort: 4317 # OTLP gRPC
- containerPort: 4318 # OTLP HTTP
- containerPort: 8889 # Prometheus metrics
volumeMounts:
- name: config
mountPath: /etc/otelcol/config.yaml
subPath: config.yaml
volumes:
- name: config
configMap:
name: otel-collector-config
---
apiVersion: v1
kind: Service
metadata:
name: otel-collector
spec:
ports:
- name: otlp-grpc
port: 4317
- name: otlp-http
port: 4318
- name: prometheus
port: 8889
selector:
app: otel-collector
EOF
Exporters
Key Concepts
Exporters send telemetry data to observability backends. OpenTelemetry supports multiple exporters simultaneously.
OTLP Exporter
The native OpenTelemetry Protocol exporter:
# Python OTLP exporter
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter as HTTPExporter
# gRPC exporter
grpc_exporter = OTLPSpanExporter(
endpoint="http://collector:4317",
insecure=True
)
# HTTP exporter
http_exporter = HTTPExporter(
endpoint="http://collector:4318/v1/traces",
headers={"Authorization": "Bearer token"}
)
Exporting to Jaeger
Jaeger (v1.35+ and all of v2) ingests OTLP natively, so the OTLP exporter
is the correct, supported way to send traces to Jaeger. The dedicated
SDK Jaeger exporters are deprecated and have been removed: the Python
opentelemetry-exporter-jaeger packages stopped being maintained in
July 2023, and the jaeger/@opentelemetry/exporter-jaeger exporters
were dropped from the SDKs. Point the OTLP exporter at Jaeger's OTLP
endpoint (port 4317 gRPC / 4318 HTTP):
# Python - send traces to Jaeger via OTLP
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
otlp_exporter = OTLPSpanExporter(endpoint="http://jaeger:4317", insecure=True)
// Node.js - send traces to Jaeger via OTLP
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const jaegerExporter = new OTLPTraceExporter({
url: 'http://jaeger:4317',
});
Prometheus Exporter
# Python Prometheus exporter
from opentelemetry.exporter.prometheus import PrometheusMetricReader
from prometheus_client import start_http_server
# Start Prometheus HTTP server
start_http_server(port=8000, addr="0.0.0.0")
# Create metric reader
reader = PrometheusMetricReader()
provider = MeterProvider(metric_readers=[reader])
// Go Prometheus exporter
import (
"go.opentelemetry.io/otel/exporters/prometheus"
"go.opentelemetry.io/otel/sdk/metric"
)
exporter, err := prometheus.New()
if err != nil {
log.Fatal(err)
}
provider := metric.NewMeterProvider(metric.WithReader(exporter))
Multiple Exporters
# Using multiple exporters simultaneously
from opentelemetry.sdk.trace.export import BatchSpanProcessor, SimpleSpanProcessor
tracer_provider = TracerProvider(resource=resource)
# Production exporter with batching
tracer_provider.add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317"))
)
# Debug exporter for development
if DEBUG:
from opentelemetry.sdk.trace.export import ConsoleSpanExporter
tracer_provider.add_span_processor(
SimpleSpanProcessor(ConsoleSpanExporter())
)
Context Propagation
Key Concepts
Context propagation carries trace context across service boundaries:
- W3C Trace Context: Standard format (
traceparent,tracestate) - B3: Zipkin format (
X-B3-TraceId,X-B3-SpanId) - Baggage: User-defined key-value pairs
Configuration
# Environment variable configuration
export OTEL_PROPAGATORS=tracecontext,baggage,b3
# Python programmatic configuration
from opentelemetry.propagate import set_global_textmap
from opentelemetry.propagators.composite import CompositePropagator
from opentelemetry.propagators.b3 import B3MultiFormat
from opentelemetry.trace.propagation.tracecontext import TraceContextTextMapPropagator
from opentelemetry.baggage.propagation import W3CBaggagePropagator
set_global_textmap(CompositePropagator([
TraceContextTextMapPropagator(),
W3CBaggagePropagator(),
B3MultiFormat(),
]))
HTTP Propagation Examples
# Python - Injecting context into outgoing requests
import requests
from opentelemetry import trace
from opentelemetry.propagate import inject
tracer = trace.get_tracer(__name__)
def call_downstream_service():
with tracer.start_as_current_span("call-service-b"):
headers = {}
inject(headers) # Inject trace context
response = requests.get(
"http://service-b/api/data",
headers=headers
)
return response.json()
# Python - Extracting context from incoming requests (Flask)
from flask import Flask, request
from opentelemetry.propagate import extract
app = Flask(__name__)
@app.route("/api/data")
def handle_request():
# Extract trace context from incoming request
context = extract(request.headers)
with tracer.start_as_current_span("handle-request", context=context):
# Process request
return {"status": "ok"}
// Node.js - Context propagation with Express
const { context, propagation, trace } = require('@opentelemetry/api');
const express = require('express');
const app = express();
// Extract context middleware
app.use((req, res, next) => {
const extractedContext = propagation.extract(context.active(), req.headers);
context.with(extractedContext, () => {
next();
});
});
// Inject context for outgoing requests
async function callService(url) {
const headers = {};
propagation.inject(context.active(), headers);
return fetch(url, { headers });
}
Baggage Usage
# Python - Using baggage
from opentelemetry import baggage
from opentelemetry.propagate import inject, extract
# Set baggage values
ctx = baggage.set_baggage("user.id", "12345")
ctx = baggage.set_baggage("tenant.id", "acme-corp", context=ctx)
# Inject into headers
headers = {}
inject(headers, context=ctx)
# Extract and read baggage
incoming_ctx = extract(request.headers)
user_id = baggage.get_baggage("user.id", context=incoming_ctx)
Common Integration Patterns
Microservices Architecture
# Service A - API Gateway
from flask import Flask
from opentelemetry.instrumentation.flask import FlaskInstrumentor
from opentelemetry.instrumentation.requests import RequestsInstrumentor
app = Flask(__name__)
FlaskInstrumentor().instrument_app(app)
RequestsInstrumentor().instrument()
@app.route("/api/orders/<order_id>")
def get_order(order_id):
with tracer.start_as_current_span("get-order") as span:
span.set_attribute("order.id", order_id)
# Call downstream services
user = requests.get(f"http://user-service/users/{user_id}").json()
products = requests.get(f"http://product-service/products").json()
return {"order": order_id, "user": user, "products": products}
Database Instrumentation
# Python - Database tracing
from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor
from sqlalchemy import create_engine
engine = create_engine("postgresql://user:pass@localhost/db")
SQLAlchemyInstrumentor().instrument(engine=engine)
# Or manual instrumentation
def query_database(query):
with tracer.start_as_current_span("db.query") as span:
span.set_attribute("db.system", "postgresql")
span.set_attribute("db.statement", query)
span.set_attribute("db.name", "mydb")
result = cursor.execute(query)
span.set_attribute("db.rows_affected", cursor.rowcount)
return result
Message Queue Integration
# Python - Kafka producer with tracing
from opentelemetry.instrumentation.kafka import KafkaInstrumentor
from kafka import KafkaProducer
KafkaInstrumentor().instrument()
producer = KafkaProducer(bootstrap_servers=['localhost:9092'])
def send_message(topic, message):
with tracer.start_as_current_span("kafka.produce") as span:
span.set_attribute("messaging.system", "kafka")
span.set_attribute("messaging.destination", topic)
# Inject context into message headers
headers = []
inject(headers, setter=kafka_header_setter)
producer.send(topic, value=message, headers=headers)
gRPC Integration
# Python - gRPC server with tracing
from opentelemetry.instrumentation.grpc import GrpcInstrumentorServer
GrpcInstrumentorServer().instrument()
class UserServiceServicer(user_pb2_grpc.UserServiceServicer):
def GetUser(self, request, context):
# Span is automatically created
current_span = trace.get_current_span()
current_span.set_attribute("user.id", request.user_id)
return user_pb2.User(id=request.user_id, name="John")
Async Operations
# Python - Async tracing
import asyncio
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
async def async_operation():
with tracer.start_as_current_span("async-parent"):
# Run multiple async operations
results = await asyncio.gather(
fetch_data("service-a"),
fetch_data("service-b"),
fetch_data("service-c"),
)
return results
async def fetch_data(service):
with tracer.start_as_current_span(f"fetch-{service}") as span:
span.set_attribute("service.name", service)
# Async HTTP call
async with aiohttp.ClientSession() as session:
async with session.get(f"http://{service}/data") as response:
return await response.json()
Quick Reference
Essential Environment Variables
| Variable | Description | Example |
|---|---|---|
OTEL_SERVICE_NAME |
Service identifier | my-service |
OTEL_RESOURCE_ATTRIBUTES |
Resource attributes | env=prod,version=1.0 |
OTEL_EXPORTER_OTLP_ENDPOINT |
Collector endpoint | http://localhost:4317 |
OTEL_EXPORTER_OTLP_PROTOCOL |
OTLP protocol | grpc or http/protobuf |
OTEL_TRACES_EXPORTER |
Trace exporter | otlp, console |
OTEL_METRICS_EXPORTER |
Metrics exporter | otlp, prometheus |
OTEL_LOGS_EXPORTER |
Logs exporter | otlp, console |
OTEL_TRACES_SAMPLER |
Sampling strategy | parentbased_traceidratio |
OTEL_TRACES_SAMPLER_ARG |
Sampler argument | 0.1 (10% sampling) |
OTEL_PROPAGATORS |
Context propagators | tracecontext,baggage,b3 |
Common Semantic Conventions
| Attribute | Description | Example |
|---|---|---|
service.name |
Service name | payment-service |
service.version |
Service version | 1.2.3 |
http.method |
HTTP method | GET, POST |
http.url |
Full URL | https://api.example.com/users |
http.status_code |
Response status | 200, 500 |
db.system |
Database type | postgresql, redis |
db.statement |
Database query | SELECT * FROM users |
messaging.system |
Message system | kafka, rabbitmq |
rpc.system |
RPC system | grpc, aws-api |
exception.type |
Exception class | ValueError |
exception.message |
Error message | Invalid input |
Collector Pipeline Components
| Component | Type | Purpose |
|---|---|---|
otlp |
Receiver | Native OTLP ingestion |
jaeger |
Receiver | Jaeger format ingestion |
prometheus |
Receiver | Prometheus scraping |
batch |
Processor | Batches telemetry |
memory_limiter |
Processor | Prevents OOM |
attributes |
Processor | Modifies attributes |
filter |
Processor | Drops unwanted data |
tail_sampling |
Processor | Intelligent sampling |
otlp |
Exporter | OTLP export (use this to reach Jaeger) |
debug |
Exporter | Console/debug output (replaces removed logging) |
prometheus |
Exporter | Prometheus exposition |
Common Issues and Solutions
Issue: Traces Not Appearing in Backend
Symptoms: Instrumented application runs but no traces visible.
Solutions:
# 1. Verify collector connectivity
curl -v http://collector:4317
# 2. Check environment variables
echo $OTEL_EXPORTER_OTLP_ENDPOINT
echo $OTEL_TRACES_EXPORTER
# 3. Enable debug logging
export OTEL_LOG_LEVEL=debug
# 4. Use console exporter for debugging
export OTEL_TRACES_EXPORTER=console
# 5. Verify SDK initialisation
from opentelemetry import trace
provider = trace.get_tracer_provider()
print(f"Provider: {provider}") # Should not be NoOpTracerProvider
Issue: High Memory Usage in Collector
Symptoms: Collector OOM or high memory consumption.
Solutions:
# Configure memory limiter processor
processors:
memory_limiter:
check_interval: 1s
limit_mib: 1000
spike_limit_mib: 200
batch:
timeout: 5s
send_batch_size: 500 # Reduce batch size
Issue: Missing Spans in Distributed Trace
Symptoms: Parent-child relationships broken across services.
Solutions:
# Ensure context propagation is configured
from opentelemetry.propagate import inject, extract
# Verify headers are injected
headers = {}
inject(headers)
print(f"Injected headers: {headers}")
# Verify extraction
context = extract(request.headers)
print(f"Extracted context: {context}")
# Check propagators are set
export OTEL_PROPAGATORS=tracecontext,baggage
Issue: Metrics Not Exported
Symptoms: Metrics created but not visible in backend.
Solutions:
# 1. Ensure metric reader is configured
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
reader = PeriodicExportingMetricReader(
exporter,
export_interval_millis=10000 # 10 seconds
)
# 2. Force flush before shutdown
meter_provider.force_flush()
meter_provider.shutdown()
Issue: High Cardinality Metrics
Symptoms: Metric explosion causing performance issues.
Solutions:
# Avoid high-cardinality attributes
# BAD: Using user_id as attribute
counter.add(1, {"user_id": user_id}) # Millions of series
# GOOD: Use bounded cardinality
counter.add(1, {"user_type": "premium"}) # Limited values
# Configure views to drop attributes
from opentelemetry.sdk.metrics import View
view = View(
instrument_name="http_requests",
attribute_keys=["method", "status"] # Only keep these
)
Issue: Sampling Not Working as Expected
Symptoms: Too many or too few traces sampled.
Solutions:
# Verify sampler configuration
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.1 # 10%
# For head-based sampling, check parent decisions
# For tail-based sampling, configure in collector
# Collector tail sampling configuration
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: sample-errors
type: status_code
status_code:
status_codes: [ERROR]
- name: sample-percentage
type: probabilistic
probabilistic:
sampling_percentage: 10
Issue: Context Lost in Async Code
Symptoms: Spans not properly nested in async operations.
Solutions:
# Python - Use context tokens
import contextvars
from opentelemetry import context
async def async_operation():
# Capture current context
ctx = context.get_current()
async def inner():
# Restore context
token = context.attach(ctx)
try:
with tracer.start_as_current_span("inner"):
pass
finally:
context.detach(token)
await inner()
// Node.js - Use context manager
const { context } = require('@opentelemetry/api');
async function asyncOperation() {
const currentContext = context.active();
await context.with(currentContext, async () => {
// Context is preserved
const span = tracer.startSpan('async-span');
// ...
span.end();
});
}
Related Topics
The following topics complement OpenTelemetry knowledge:
- Jaeger - Distributed tracing backend that works seamlessly with OpenTelemetry, offering advanced trace visualisation and analysis
- Prometheus - Metrics collection and alerting system commonly used alongside OpenTelemetry for comprehensive monitoring
- Grafana - Visualisation platform for creating dashboards from OpenTelemetry metrics and traces
- Kubernetes Observability - Deploying and configuring OpenTelemetry in Kubernetes environments with operators and sidecars
- Service Mesh Integration - Using OpenTelemetry with Istio, Linkerd, or other service meshes for enhanced observability
- Log Aggregation (ELK/Loki) - Correlating OpenTelemetry logs with traces for complete observability