Datadog
Cloud-based monitoring and analytics platform for infrastructure, applications, logs, and traces with unified observability.
Datadog
Cloud-based monitoring and analytics platform for infrastructure, applications, logs, and traces with unified observability.
Overview
Datadog is a SaaS-based monitoring and analytics platform that provides full-stack observability across infrastructure, applications, logs, and traces. It uses an agent-based approach to collect metrics, logs, and traces from hosts and containers, sending data to Datadog's cloud platform for analysis, visualisation, and alerting. The platform integrates with hundreds of technologies and supports custom instrumentation.
graph TB
subgraph "Data Sources"
A[Hosts/VMs]
B[Containers]
C[Cloud Services]
D[Applications]
end
subgraph "Collection Layer"
E[Datadog Agent]
F[DogStatsD]
G[API/SDK]
H[Cloud Integrations]
end
subgraph "Datadog Platform"
I[Metrics Storage]
J[Log Storage]
K[APM Storage]
L[Query Engine]
M[Dashboards]
N[Monitors/Alerts]
O[Analytics]
end
A --> E
B --> E
C --> H
D --> G
E --> F
E --> I
E --> J
F --> I
G --> K
H --> I
I --> L
J --> L
K --> L
L --> M
L --> N
L --> O
Agent Installation and Configuration
The Datadog Agent is a lightweight daemon that collects metrics, logs, and traces from your hosts.
Installation
# Ubuntu/Debian, RedHat/CentOS/Amazon Linux (one script for all distros)
# The legacy s3.amazonaws.com/dd-agent install_script.sh is deprecated and
# defaults to Agent v6; use the versioned install.datadoghq.com script.
DD_API_KEY=<YOUR_API_KEY> DD_SITE="datadoghq.com" \
bash -c "$(curl -L https://install.datadoghq.com/scripts/install_script_agent7.sh)"
# Docker
docker run -d --name datadog-agent \
-e DD_API_KEY=<YOUR_API_KEY> \
-e DD_SITE="datadoghq.com" \
-v /var/run/docker.sock:/var/run/docker.sock:ro \
-v /proc/:/host/proc/:ro \
-v /sys/fs/cgroup/:/host/sys/fs/cgroup:ro \
gcr.io/datadoghq/agent:7
# Kubernetes (Helm)
helm repo add datadog https://helm.datadoghq.com
helm install datadog-agent datadog/datadog \
--set datadog.apiKey=<YOUR_API_KEY> \
--set datadog.site=datadoghq.com
Main Configuration File
# /etc/datadog-agent/datadog.yaml
api_key: <YOUR_API_KEY>
site: datadoghq.com
# Hostname (optional, auto-detected by default)
hostname: web-server-01
# Tags for this host
tags:
- env:production
- role:webserver
- team:platform
# Enable logs collection
logs_enabled: true
# Enable APM
apm_config:
enabled: true
env: production
# Enable process monitoring
process_config:
enabled: true
# Network performance monitoring
network_config:
enabled: true
# Enable live container monitoring
container_collect_all: true
Agent Commands
# Start/stop/restart agent
sudo systemctl start datadog-agent
sudo systemctl stop datadog-agent
sudo systemctl restart datadog-agent
# Check agent status
sudo datadog-agent status
# Check specific integration status
sudo datadog-agent status integrations
# Validate configuration
sudo datadog-agent configcheck
# Test connectivity
sudo datadog-agent diagnose
# Restart specific integration
sudo datadog-agent restart <integration_name>
Docker Agent Configuration
# docker-compose.yml
version: '3'
services:
datadog-agent:
image: gcr.io/datadoghq/agent:7
environment:
- DD_API_KEY=${DD_API_KEY}
- DD_SITE=datadoghq.com
- DD_LOGS_ENABLED=true
- DD_APM_ENABLED=true
- DD_APM_NON_LOCAL_TRAFFIC=true
- DD_PROCESS_AGENT_ENABLED=true
- DD_TAGS=env:production role:docker
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
- /proc/:/host/proc/:ro
- /sys/fs/cgroup:/host/sys/fs/cgroup:ro
- /etc/datadog-agent/conf.d:/etc/datadog-agent/conf.d:ro
ports:
- "8126:8126/tcp" # APM
- "8125:8125/udp" # DogStatsD
Metrics Collection
Datadog collects standard infrastructure metrics automatically and supports custom metrics through various methods.
Standard Metrics
# System metrics (collected automatically)
system.cpu.user # CPU usage by user processes
system.cpu.system # CPU usage by system processes
system.mem.used # Memory used
system.mem.free # Memory free
system.disk.used # Disk space used
system.disk.free # Disk space free
system.net.bytes_sent # Network bytes sent
system.net.bytes_rcvd # Network bytes received
system.load.1 # 1-minute load average
system.load.5 # 5-minute load average
system.load.15 # 15-minute load average
# Docker metrics (when Docker integration enabled)
docker.containers.running # Number of running containers
docker.containers.stopped # Number of stopped containers
docker.cpu.usage # Container CPU usage
docker.mem.usage # Container memory usage
Custom Metrics with DogStatsD
# Python - using datadog library
from datadog import initialize, statsd
options = {
'statsd_host': '127.0.0.1',
'statsd_port': 8125
}
initialize(**options)
# Counter - tracks cumulative count
statsd.increment('page.views', tags=['page:home', 'env:prod'])
statsd.increment('api.requests', value=5)
# Gauge - tracks current value
statsd.gauge('database.connections', 42, tags=['db:mysql'])
statsd.gauge('queue.size', 100)
# Histogram - statistical distribution
statsd.histogram('request.duration', 0.235, tags=['endpoint:/api/users'])
# Distribution - global percentiles
statsd.distribution('request.size', 1024, tags=['method:POST'])
# Set - count unique values
statsd.set('unique.visitors', 'user123', tags=['site:main'])
# Timing - convenience for histograms
statsd.timing('query.time', 350) # milliseconds
# Context manager for timing
with statsd.timed('database.query', tags=['query:users']):
# Your code here
pass
// Node.js - using hot-shots library
const StatsD = require('hot-shots');
const dogstatsd = new StatsD({
host: 'localhost',
port: 8125,
globalTags: ['env:prod']
});
// Counter
dogstatsd.increment('page.views', 1, ['page:home']);
// Gauge
dogstatsd.gauge('active.connections', 42);
// Histogram
dogstatsd.histogram('response.time', 235, ['endpoint:/api']);
// Distribution
dogstatsd.distribution('file.size', 2048);
// Timing
dogstatsd.timing('query.duration', 150);
// Set
dogstatsd.set('unique.users', 'user456');
// Go - using datadog-go library
import "github.com/DataDog/datadog-go/v5/statsd"
client, err := statsd.New("127.0.0.1:8125",
statsd.WithNamespace("myapp."),
statsd.WithTags([]string{"env:prod"}),
)
if err != nil {
log.Fatal(err)
}
defer client.Close()
// Counter
client.Incr("page.views", []string{"page:home"}, 1)
// Gauge
client.Gauge("queue.size", 100, []string{"queue:jobs"}, 1)
// Histogram
client.Histogram("request.duration", 0.235, []string{"endpoint:/api"}, 1)
// Distribution
client.Distribution("request.size", 1024, []string{"method:POST"}, 1)
// Timing
client.Timing("query.time", time.Millisecond*350, []string{"db:postgres"}, 1)
Custom Metrics via Agent Checks
# /etc/datadog-agent/checks.d/custom_check.py
from datadog_checks.base import AgentCheck
class CustomCheck(AgentCheck):
def check(self, instance):
# Gauge metric
self.gauge(
'custom.metric.value',
42,
tags=['env:prod', 'service:api']
)
# Counter metric
self.count(
'custom.events.total',
1,
tags=['type:login']
)
# Rate metric (count per second)
self.rate(
'custom.requests.rate',
10
)
# Service check
self.service_check(
'custom.service.status',
AgentCheck.OK,
message='Service is healthy',
tags=['service:api']
)
# /etc/datadog-agent/conf.d/custom_check.yaml
init_config:
instances:
- min_collection_interval: 60 # Run every 60 seconds
Log Collection and Parsing
Datadog Agent can collect logs from various sources and parse them for analysis.
Enable Log Collection
# /etc/datadog-agent/datadog.yaml
logs_enabled: true
# Optional log processing
logs_config:
# Process raw logs on agent before sending
use_http: true
use_compression: true
compression_level: 6
# Additional processing
processing_rules:
- type: exclude_at_match
name: exclude_healthcheck
pattern: /healthcheck
File-based Log Collection
# /etc/datadog-agent/conf.d/custom_logs.yaml
logs:
- type: file
path: /var/log/myapp/*.log
service: myapp
source: python
sourcecategory: sourcecode
tags:
- env:production
- team:backend
# Multi-line log aggregation
log_processing_rules:
- type: multi_line
name: log_start_pattern
pattern: \d{4}-\d{2}-\d{2} # Lines starting with date
# JSON logs
- type: file
path: /var/log/myapp/json.log
service: myapp
source: python
sourcecategory: sourcecode
# Nginx access logs
- type: file
path: /var/log/nginx/access.log
service: nginx
source: nginx
sourcecategory: http_web_access
Docker Log Collection
# docker-compose.yml labels
services:
web:
image: myapp:latest
labels:
com.datadoghq.ad.logs: '[{"source": "python", "service": "myapp"}]'
com.datadoghq.ad.tags: '["env:prod", "version:1.2.3"]'
# Kubernetes pod annotations
apiVersion: v1
kind: Pod
metadata:
name: myapp
annotations:
ad.datadoghq.com/myapp.logs: |
[
{
"source": "python",
"service": "myapp",
"tags": ["env:prod"]
}
]
spec:
containers:
- name: myapp
image: myapp:latest
Log Processing Rules
# /etc/datadog-agent/conf.d/custom_logs.yaml
logs:
- type: file
path: /var/log/myapp/app.log
service: myapp
source: python
log_processing_rules:
# Exclude logs matching pattern
- type: exclude_at_match
name: exclude_debug
pattern: DEBUG
# Include only logs matching pattern
- type: include_at_match
name: include_errors
pattern: ERROR|CRITICAL
# Mask sensitive data
- type: mask_sequences
name: mask_credit_cards
pattern: \d{4}-\d{4}-\d{4}-\d{4}
replace_placeholder: "[MASKED_CC]"
# Multi-line aggregation
- type: multi_line
name: aggregate_stack_traces
pattern: ^\s+at
Log Pipeline and Parsing
// Grok parser example (configured in Datadog UI)
// Pattern for parsing custom log format
rule %{date("yyyy-MM-dd HH:mm:ss"):timestamp} \[%{word:log_level}\] %{data:logger} - %{data:message}
// Example log:
// 2025-12-03 10:30:45 [ERROR] app.handlers - Database connection failed
// Parsed attributes:
// timestamp: 2025-12-03 10:30:45
// log_level: ERROR
// logger: app.handlers
// message: Database connection failed
Common Log Patterns
# Application logs
logs:
- type: file
path: /var/log/app/application.log
service: myapp
source: java
log_processing_rules:
- type: multi_line
name: java_stack_trace
pattern: ^\s+(at|\.{3})\s+
# Syslog
logs:
- type: tcp
port: 10514
service: syslog
source: syslog
# JSON logs (auto-parsed)
logs:
- type: file
path: /var/log/app/json.log
service: myapp
source: nodejs
# JSON logs are automatically parsed
APM and Distributed Tracing
Application Performance Monitoring tracks requests across services with distributed tracing.
sequenceDiagram
participant Client
participant WebService
participant APIService
participant Database
participant DatadogAgent
Client->>WebService: HTTP Request
WebService->>APIService: Internal API Call
APIService->>Database: Query
Database-->>APIService: Result
APIService-->>WebService: Response
WebService-->>Client: HTTP Response
WebService->>DatadogAgent: Trace Span
APIService->>DatadogAgent: Trace Span
DatadogAgent->>Datadog: Traces
Enable APM Agent
# /etc/datadog-agent/datadog.yaml
apm_config:
enabled: true
env: production
# Allow traces from other containers
apm_non_local_traffic: true
# Ingestion sampling (current mechanism; the legacy App Analytics
# `analyzed_rate_by_service` is superseded).
# Head-based per-service/resource sampling rules in the Agent:
trace_sampling_rules:
- service: my-service
sample_rate: 1.0 # Keep 100% of this service's traces
# Per-host throughput cap for the Agent's automatic sampler (env: DD_APM_MAX_TPS):
max_traces_per_second: 50
# Tracing-library-side rules use DD_TRACE_SAMPLING_RULES instead.
# For volume management, also see Datadog ingestion controls + retention filters.
# Resource filtering
filter_tags:
require: ["env:production"]
# Obfuscation
obfuscation:
elasticsearch:
enabled: true
mongodb:
enabled: true
http:
remove_query_string: true
remove_paths_with_digits: true
redis:
enabled: true
Python APM Instrumentation
# Automatic instrumentation
from ddtrace import patch_all
patch_all()
# Application code
from flask import Flask
app = Flask(__name__)
@app.route('/')
def hello():
return 'Hello World!'
if __name__ == '__main__':
app.run()
# Run with ddtrace
# DD_SERVICE=myapp DD_ENV=prod DD_VERSION=1.0 \
# ddtrace-run python app.py
# Manual instrumentation
from ddtrace import tracer
@tracer.wrap(service='myapp', resource='process_data')
def process_data(data):
with tracer.trace('database.query', service='postgres') as span:
span.set_tag('query.type', 'SELECT')
span.set_tag('rows.count', 100)
result = db.query(data)
return result
# Custom span
with tracer.trace('custom.operation') as span:
span.set_tag('user.id', user_id)
span.set_metric('items.processed', 42)
# Your code here
Node.js APM Instrumentation
// tracer.js - initialize first
const tracer = require('dd-trace').init({
service: 'myapp',
env: 'production',
version: '1.0.0',
analytics: true,
logInjection: true
});
module.exports = tracer;
// app.js
require('./tracer'); // Must be first
const express = require('express');
const app = express();
// Automatic instrumentation for supported libraries
app.get('/', (req, res) => {
res.send('Hello World!');
});
// Manual instrumentation
const tracer = require('dd-trace');
function processOrder(order) {
const span = tracer.startSpan('process.order', {
tags: {
'order.id': order.id,
'order.amount': order.amount
}
});
try {
// Your code here
span.setTag('result', 'success');
} catch (err) {
span.setTag('error', true);
span.setTag('error.msg', err.message);
throw err;
} finally {
span.finish();
}
}
Go APM Instrumentation
package main
import (
"net/http"
httptrace "gopkg.in/DataDog/dd-trace-go.v1/contrib/net/http"
"gopkg.in/DataDog/dd-trace-go.v1/ddtrace/tracer"
)
func main() {
// Initialize tracer
tracer.Start(
tracer.WithService("myapp"),
tracer.WithEnv("production"),
tracer.WithServiceVersion("1.0.0"),
)
defer tracer.Stop()
// Automatic HTTP instrumentation
mux := httptrace.NewServeMux()
mux.HandleFunc("/", handler)
http.ListenAndServe(":8080", mux)
}
func handler(w http.ResponseWriter, r *http.Request) {
// Manual span creation
span, ctx := tracer.StartSpanFromContext(r.Context(), "custom.operation")
defer span.Finish()
span.SetTag("user.id", "123")
// Your code here
processData(ctx)
}
func processData(ctx context.Context) {
span, _ := tracer.StartSpanFromContext(ctx, "process.data")
defer span.Finish()
// Your code here
}
Java APM Instrumentation
# Download Java agent
wget -O dd-java-agent.jar \
https://dtdg.co/latest-java-tracer
# Run application with agent
java -javaagent:dd-java-agent.jar \
-Ddd.service=myapp \
-Ddd.env=production \
-Ddd.version=1.0.0 \
-Ddd.trace.analytics.enabled=true \
-jar myapp.jar
Service Map Configuration
# Tag spans for service mapping
# Tags are automatically detected from:
# - service: service name
# - env: environment
# - version: application version
# - resource: specific endpoint or operation
# Example span tags
span.service = "web-frontend"
span.env = "production"
span.version = "2.1.0"
span.resource = "GET /api/users"
span.type = "web"
Dashboard Creation
Datadog dashboards provide customisable visualisations for metrics, logs, and traces.
Dashboard Types
| Type | Description | Use Case |
|---|---|---|
| Timeboard | Fixed time synchronisation across all widgets | System monitoring, correlation analysis |
| Screenboard | Free-form layout with independent time ranges | Executive dashboards, status boards |
Common Widget Types
// Timeseries widget
{
"definition": {
"type": "timeseries",
"requests": [
{
"q": "avg:system.cpu.user{*}",
"display_type": "line",
"style": {
"palette": "dog_classic",
"line_type": "solid",
"line_width": "normal"
}
}
],
"title": "CPU Usage",
"show_legend": true,
"legend_size": "0"
}
}
// Query value widget (single metric)
{
"definition": {
"type": "query_value",
"requests": [
{
"q": "avg:system.mem.used{*}",
"aggregator": "avg"
}
],
"title": "Memory Used",
"precision": 2,
"autoscale": true
}
}
// Heatmap widget
{
"definition": {
"type": "heatmap",
"requests": [
{
"q": "avg:trace.http.request.duration{*} by {resource_name}"
}
],
"title": "Request Duration Heatmap"
}
}
// Top list widget
{
"definition": {
"type": "toplist",
"requests": [
{
"q": "top(avg:docker.cpu.usage{*} by {container_name}, 10, 'mean', 'desc')"
}
],
"title": "Top 10 Containers by CPU"
}
}
Dashboard via API
from datadog import initialize, api
initialize(api_key='<YOUR_API_KEY>', app_key='<YOUR_APP_KEY>')
# Create dashboard
dashboard = {
'title': 'System Overview',
'description': 'System metrics dashboard',
'layout_type': 'ordered',
'widgets': [
{
'definition': {
'type': 'timeseries',
'requests': [
{
'q': 'avg:system.cpu.user{*}',
'display_type': 'line'
}
],
'title': 'CPU Usage'
}
},
{
'definition': {
'type': 'query_value',
'requests': [
{
'q': 'avg:system.mem.used{*}',
'aggregator': 'last'
}
],
'title': 'Memory Used',
'autoscale': True
}
}
]
}
result = api.Dashboard.create(**dashboard)
print(f"Dashboard created: {result['id']}")
Template Variables
// Dashboard with template variables
{
"title": "Infrastructure Dashboard",
"template_variables": [
{
"name": "env",
"default": "production",
"prefix": "env"
},
{
"name": "host",
"default": "*",
"prefix": "host"
},
{
"name": "service",
"default": "*",
"prefix": "service"
}
],
"widgets": [
{
"definition": {
"type": "timeseries",
"requests": [
{
"q": "avg:system.cpu.user{$env,$host,$service}"
}
]
}
}
]
}
Monitor and Alert Configuration
Monitors detect conditions and trigger alerts based on metric thresholds and anomalies.
flowchart TD
A[Metric Data] --> B{Monitor Evaluation}
B -->|Threshold Exceeded| C[Alert State]
B -->|Within Bounds| D[OK State]
C --> E{Notification Channels}
E --> F[Email]
E --> G[Slack]
E --> H[PagerDuty]
E --> I[Webhook]
D --> J[No Action]
C --> K[Alert Recovery]
K -->|Metric Normal| D
Monitor Types
| Type | Description | Use Case |
|---|---|---|
| Metric | Alert on metric threshold | CPU > 80%, memory usage |
| APM | Alert on trace metrics | Error rate, latency p99 |
| Integration | Cloud service monitoring | AWS RDS, ELB status |
| Process | Monitor process availability | Nginx running, pod count |
| Network | Network performance | TCP connections, bandwidth |
| Log | Alert on log patterns | Error log frequency |
| Event | Alert on custom events | Deployment events |
| Composite | Combine multiple monitors | Complex conditions |
| Anomaly | ML-based anomaly detection | Unusual traffic patterns |
| Outlier | Detect outlier hosts | One host behaving differently |
| Forecast | Predict future values | Disk will fill in 2 days |
Metric Monitor
// Simple threshold monitor
{
"name": "High CPU Usage",
"type": "metric alert",
"query": "avg(last_5m):avg:system.cpu.user{env:production} > 80",
"message": "CPU usage is above 80% on {{host.name}} @slack-alerts",
"tags": ["team:infrastructure", "priority:high"],
"options": {
"notify_audit": false,
"locked": false,
"timeout_h": 0,
"include_tags": true,
"no_data_timeframe": 10,
"require_full_window": false,
"new_host_delay": 300,
"notify_no_data": true,
"renotify_interval": 0,
"escalation_message": "CPU still high after 30 minutes",
"thresholds": {
"critical": 80,
"warning": 70,
"critical_recovery": 75,
"warning_recovery": 65
}
}
}
Multi-Alert Monitor
// Alert per service
{
"name": "High Error Rate per Service",
"type": "metric alert",
"query": "avg(last_10m):sum:trace.http.request.errors{env:production} by {service}.as_rate() > 5",
"message": "Error rate for {{service.name}} is above 5% @pagerduty",
"options": {
"thresholds": {
"critical": 5,
"warning": 2
},
"notify_no_data": false,
"no_data_timeframe": 20
}
}
Anomaly Detection Monitor
// Anomaly detection for traffic patterns
{
"name": "Anomalous Request Rate",
"type": "query alert",
"query": "avg(last_4h):anomalies(avg:trace.http.request.hits{env:production}.as_rate(), 'basic', 2, direction='both', alert_window='last_15m', interval=60, count_default_zero='true') >= 1",
"message": "Unusual request rate detected @slack-alerts",
"options": {
"thresholds": {
"critical": 1,
"critical_recovery": 0
},
"threshold_windows": {
"trigger_window": "last_15m",
"recovery_window": "last_15m"
}
}
}
Log Monitor
// Alert on error log frequency
{
"name": "High Error Log Rate",
"type": "log alert",
"query": "logs(\"status:error service:myapp\").index(\"main\").rollup(\"count\").last(\"5m\") > 100",
"message": "More than 100 errors in last 5 minutes @slack-alerts\n{{#is_alert}}\nError details: {{log.message}}\n{{/is_alert}}",
"options": {
"thresholds": {
"critical": 100,
"warning": 50
},
"enable_logs_sample": true
}
}
Composite Monitor
# Create composite monitor via API
from datadog import initialize, api
initialize(api_key='<API_KEY>', app_key='<APP_KEY>')
# Composite monitor combining CPU and memory
monitor = {
'name': 'High Resource Usage',
'type': 'composite',
'query': '(monitor_id_1 && monitor_id_2) || monitor_id_3',
'message': 'Critical resource usage detected @pagerduty',
'options': {
'notify_no_data': False
}
}
api.Monitor.create(**monitor)
Monitor Templates
// Monitor with template variables
{
"name": "{{service.name}} - High Latency",
"type": "metric alert",
"query": "avg(last_5m):avg:trace.http.request.duration.by.service.99p{service:{{service.name}},env:production} > 1",
"message": "P99 latency for {{service.name}} is above 1s\n\nHost: {{host.name}}\nEnv: {{env.name}}\n\n@slack-{{service.name}}",
"tags": ["service:{{service.name}}", "auto-generated"]
}
Notification Channels
# Alert message format with @ mentions
@slack-alerts # Slack channel
@pagerduty-critical # PagerDuty service
@webhook-custom # Custom webhook
@user@example.com # Email address
@all # All team members
# Conditional notifications
{{#is_alert}}
Alert triggered!
{{/is_alert}}
{{#is_warning}}
Warning level reached
{{/is_warning}}
{{#is_recovery}}
Issue resolved
{{/is_recovery}}
{{#is_no_data}}
No data received
{{/is_no_data}}
Monitor via API
from datadog import initialize, api
initialize(api_key='<API_KEY>', app_key='<APP_KEY>')
# Create monitor
monitor = {
'name': 'Disk Space Low',
'type': 'metric alert',
'query': 'avg(last_5m):avg:system.disk.free{*} by {host} < 10000000000',
'message': '''Disk space below 10GB on {{host.name}}
Current value: {{value}}
Threshold: {{threshold}}
@slack-infrastructure''',
'tags': ['auto:true', 'team:platform'],
'options': {
'thresholds': {
'critical': 10000000000,
'warning': 20000000000
},
'notify_no_data': True,
'no_data_timeframe': 20,
'require_full_window': False,
'notify_audit': False,
'include_tags': True
}
}
result = api.Monitor.create(**monitor)
print(f"Monitor created: {result['id']}")
# Update monitor
api.Monitor.update(
monitor_id,
query='avg(last_5m):avg:system.disk.free{*} by {host} < 5000000000'
)
# Delete monitor
api.Monitor.delete(monitor_id)
# Mute monitor
api.Monitor.mute(monitor_id, end=1735948800) # Unix timestamp
# Unmute monitor
api.Monitor.unmute(monitor_id)
Common Integrations
Datadog integrates with hundreds of services and platforms for automatic metric collection.
AWS Integration
# Configure via Datadog UI or API
# Requires IAM role with permissions:
# - cloudwatch:GetMetricStatistics
# - cloudwatch:ListMetrics
# - ec2:DescribeInstances
# - ec2:DescribeTags
# - rds:DescribeDBInstances
# - s3:ListBucket
# Automatically collects:
# - EC2 instance metrics
# - RDS database metrics
# - ELB load balancer metrics
# - Lambda function metrics
# - S3 bucket metrics
# - CloudWatch custom metrics
Kubernetes Integration
# values.yaml for Helm chart
datadog:
apiKey: <YOUR_API_KEY>
site: datadoghq.com
# Cluster monitoring
clusterName: production-cluster
# Collect Kubernetes events
collectEvents: true
# Leader election for cluster checks
leaderElection: true
# Kubernetes state metrics
kubeStateMetricsEnabled: true
# APM
apm:
enabled: true
port: 8126
# Logs
logs:
enabled: true
containerCollectAll: true
# Process monitoring
processAgent:
enabled: true
# Network performance monitoring
networkMonitoring:
enabled: true
# Cluster checks
clusterChecks:
enabled: true
# Kubernetes metrics collected automatically:
# - kubernetes.cpu.usage
# - kubernetes.memory.usage
# - kubernetes.network.rx_bytes
# - kubernetes.network.tx_bytes
# - kubernetes_state.deployment.replicas_desired
# - kubernetes_state.pod.ready
PostgreSQL Integration
# /etc/datadog-agent/conf.d/postgres.d/conf.yaml
init_config:
instances:
- host: localhost
port: 5432
username: datadog
password: '<PASSWORD>'
dbname: postgres
# Collect additional metrics
collect_function_metrics: true
collect_count_metrics: true
collect_activity_metrics: true
collect_database_size_metrics: true
collect_default_database: true
# Custom queries
custom_queries:
- metric_prefix: postgresql
query: SELECT count(*) as connections FROM pg_stat_activity;
columns:
- name: connections
type: gauge
tags:
- custom_query:active_connections
tags:
- env:production
- db:main
-- Create Datadog user in PostgreSQL
CREATE USER datadog WITH PASSWORD '<PASSWORD>';
GRANT pg_monitor TO datadog;
GRANT SELECT ON pg_stat_database TO datadog;
Redis Integration
# /etc/datadog-agent/conf.d/redisdb.d/conf.yaml
init_config:
instances:
- host: localhost
port: 6379
password: '<PASSWORD>'
# Collect slow commands
command_stats: true
# Warn on slow operations
slowlog-max-len: 128
tags:
- env:production
- cache:main
NGINX Integration
# /etc/datadog-agent/conf.d/nginx.d/conf.yaml
init_config:
instances:
- nginx_status_url: http://localhost/nginx_status/
# Additional configurations
tags:
- env:production
- instance:web-1
# Enable NGINX stub_status module
# /etc/nginx/sites-available/default
server {
listen 80;
location /nginx_status {
stub_status on;
access_log off;
allow 127.0.0.1;
deny all;
}
}
MySQL Integration
# /etc/datadog-agent/conf.d/mysql.d/conf.yaml
init_config:
instances:
- host: localhost
port: 3306
user: datadog
pass: '<PASSWORD>'
# Replication metrics
replication: true
# InnoDB metrics
extra_innodb_metrics: true
# Performance schema
extra_performance_metrics: true
# Schema metrics
schema_size_metrics: true
tags:
- env:production
- db:primary
-- Create Datadog user in MySQL
CREATE USER 'datadog'@'localhost' IDENTIFIED BY '<PASSWORD>';
GRANT REPLICATION CLIENT ON *.* TO 'datadog'@'localhost';
GRANT PROCESS ON *.* TO 'datadog'@'localhost';
GRANT SELECT ON performance_schema.* TO 'datadog'@'localhost';
Docker Integration
# Automatically enabled when Docker socket is mounted
# Collects container metrics:
# - docker.cpu.usage
# - docker.mem.usage
# - docker.io.read_bytes
# - docker.io.write_bytes
# - docker.net.bytes_sent
# - docker.net.bytes_rcvd
# docker-compose.yml
services:
datadog-agent:
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
environment:
- DD_CONTAINER_EXCLUDE: "name:datadog-agent" # Exclude self
- DD_CONTAINER_INCLUDE: "image:myapp.*" # Include specific images
Query Language Basics
Datadog uses a query language for metrics, logs, and traces.
Metric Query Syntax
# Basic structure
function(aggregation:metric_name{tag_filters}[.rollup(method, time)])
# Examples
avg:system.cpu.user{host:web-01}
sum:http.requests{service:api,env:prod}.as_count()
max:system.mem.used{*} by {host}
p95:trace.http.request.duration{service:web}
# Aggregation functions
avg # Average
sum # Sum
min # Minimum
max # Maximum
count # Count
# Time aggregation (rollup)
.as_count() # Sum over rollup period
.as_rate() # Per-second rate
.rollup(avg, 60) # Average over 60 seconds
Metric Functions
# Arithmetic
avg:system.cpu.user{*} + avg:system.cpu.system{*}
avg:system.mem.used{*} / avg:system.mem.total{*} * 100
# Rate of change
per_second(sum:http.requests{*}.as_count())
per_minute(sum:errors.total{*})
# Derivative (change between consecutive points)
derivative(avg:system.disk.used{*})
# Moving average
ewma_10(avg:system.load.1{*}) # Exponentially weighted moving average
autosmooth(avg:latency{*}) # Automatic smoothing
# Anomaly detection
anomalies(avg:requests.count{*}, 'basic', 2)
anomalies(avg:cpu.usage{*}, 'agile', 3, direction='above')
# Forecasting
forecast(avg:disk.used{*}, 'linear', 2) # Forecast 2 periods ahead
# Rollup functions
.rollup(avg, 300) # 5-minute average
.rollup(sum, 3600) # 1-hour sum
.rollup(max, 60) # 1-minute maximum
# Top/bottom
top(avg:cpu.usage{*} by {host}, 5, 'mean', 'desc')
bottom(avg:latency{*} by {endpoint}, 3, 'last', 'asc')
Tag Filtering
# Exact match
metric{tag:value}
metric{env:production}
# Multiple tags (AND)
metric{env:production,service:api}
# Wildcard
metric{host:web-*}
metric{service:*api*}
# OR condition
metric{env:production OR env:staging}
# NOT condition
metric{env:production,host:!web-01}
# Group by tags
avg:metric{*} by {host}
sum:metric{*} by {service,env}
Log Query Syntax
# Basic search
status:error
service:myapp
message:"connection timeout"
# Boolean operators
status:error AND service:myapp
status:(error OR warning)
status:error AND NOT service:healthcheck
# Wildcards
service:web-*
message:*timeout*
# Numeric ranges
http.status_code:[400 TO 499]
duration:>1000
response_time:[100 TO 500]
# Facet search
@http.method:POST
@user.id:12345
@error.kind:DatabaseError
# Nested fields (JSON)
@attributes.user.email:"user@example.com"
# Existence check
@http.status_code:* # Has status code field
-@http.status_code:* # Missing status code field
Log Aggregations
# Count
count() by @service
# Measure
avg(@duration) by @http.method
max(@response_time) by @endpoint
p95(@latency) by @region
# Multiple dimensions
count() by @service,@env
avg(@duration) by @service,@http.status_code
APM Query Syntax
# Trace search
service:web-frontend
resource:"GET /api/users"
env:production
# Error traces
error:true
@http.status_code:[500 TO 599]
# Duration filtering
duration:>1s
@duration:[100ms TO 500ms]
# Resource patterns
resource:"GET /api/*"
operation:http.request
# Span tags
@http.method:POST
@db.type:postgres
@user.id:12345
# Analytics
avg(@duration) by service
p99(@duration) by resource
count() by @http.status_code
Time Range Functions
# Relative time ranges
last_5m # Last 5 minutes
last_1h # Last 1 hour
last_4h # Last 4 hours
last_1d # Last 1 day
last_1w # Last 1 week
# Compare to previous period
avg:metric{*}, avg:metric{*}.shift(-1w) # Compare to last week
avg:metric{*}, avg:metric{*}.shift(-1d) # Compare to yesterday
# Time shift
.shift(-1h) # Shift back 1 hour
.shift(1d) # Shift forward 1 day
Quick Reference
Essential Commands
# Agent management
sudo systemctl start datadog-agent
sudo systemctl stop datadog-agent
sudo systemctl restart datadog-agent
sudo datadog-agent status
sudo datadog-agent configcheck
# Testing
sudo datadog-agent check <integration_name>
sudo datadog-agent diagnose
sudo -u dd-agent datadog-agent check <check_name>
# Logs
sudo tail -f /var/log/datadog/agent.log
sudo journalctl -u datadog-agent -f
Common Metrics
# System
system.cpu.user # CPU user time
system.cpu.system # CPU system time
system.mem.used # Memory used
system.disk.used # Disk used
system.net.bytes_sent # Network sent
system.load.1 # 1-min load average
# Docker
docker.cpu.usage # Container CPU
docker.mem.usage # Container memory
docker.containers.running # Running containers
# APM
trace.http.request.hits # Request count
trace.http.request.errors # Error count
trace.http.request.duration # Request duration
Key Configuration Files
# Main configuration
/etc/datadog-agent/datadog.yaml
# Integration configs
/etc/datadog-agent/conf.d/<integration>.d/conf.yaml
# Custom checks
/etc/datadog-agent/checks.d/
/etc/datadog-agent/conf.d/
# Logs
/var/log/datadog/agent.log
/var/log/datadog/trace-agent.log
API Example
from datadog import initialize, api
initialize(api_key='<API_KEY>', app_key='<APP_KEY>')
# Post metric
api.Metric.send(
metric='custom.metric',
points=[(time.time(), 42)],
tags=['env:prod']
)
# Query metrics
api.Metric.query(
start=int(time.time()) - 3600,
end=int(time.time()),
query='avg:system.cpu.user{*}'
)
# Create monitor
api.Monitor.create(
type='metric alert',
query='avg(last_5m):avg:system.cpu.user{*} > 80',
name='High CPU',
message='@slack-alerts'
)
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Agent not reporting | API key incorrect | Verify API key in datadog.yaml, check datadog-agent status |
| High agent CPU usage | Too many integrations/checks | Reduce check frequency, disable unused integrations |
| Missing metrics | Integration not configured | Check integration config in /etc/datadog-agent/conf.d/ |
| No logs appearing | Logs not enabled | Set logs_enabled: true in datadog.yaml |
| APM traces missing | Agent APM disabled | Enable with apm_config.enabled: true |
| Service not traced | Library not instrumented | Install and configure APM library for language |
| High data usage | Too many custom metrics | Use metric aggregation, reduce cardinality, filter tags |
| Duplicate metrics | Multiple agents reporting | Check hostname configuration, ensure unique hostnames |
| Monitor flapping | Threshold too sensitive | Adjust threshold or increase evaluation window |
| No data in dashboards | Wrong time range selected | Verify time range and template variable selections |
| Container metrics missing | Docker socket not mounted | Mount /var/run/docker.sock in agent container |
| Kubernetes metrics missing | Agent not deployed as DaemonSet | Deploy agent on every node, enable cluster checks |
| Slow dashboard loading | Too many queries | Reduce number of widgets, increase time granularity |
| API rate limit exceeded | Too many API calls | Implement rate limiting, batch requests |
| Logs not parsed | Missing pipeline | Create log pipeline with parsers in Datadog UI |
| High latency in traces | Sampling rate too high | Reduce sampling rate in APM config |
Debugging Steps
# Check agent status
sudo datadog-agent status
# Test specific integration
sudo datadog-agent check <integration_name> -l debug
# Check connectivity
sudo datadog-agent diagnose
# View recent logs
sudo tail -f /var/log/datadog/agent.log
# Check configuration validity
sudo datadog-agent configcheck
# List running checks
sudo datadog-agent check list
# Flare (send diagnostics to support)
sudo datadog-agent flare <case-id>
Performance Optimisation
# /etc/datadog-agent/datadog.yaml
# Reduce collection frequency
check_runners: 4
# Increase batch size for metrics
aggregator_buffer_size: 100
# Limit container collection
container_exclude: ["name:.*-test", "image:.*debug.*"]
container_include: ["name:prod-.*"]
# Reduce log collection
logs_config:
# Increase batch size
batch_wait: 5
# Compress logs
use_compression: true
compression_level: 6
Custom Check Troubleshooting
# Test custom check
sudo -u dd-agent datadog-agent check <check_name> -l debug
# Check Python syntax
python3 -m py_compile /etc/datadog-agent/checks.d/<check_name>.py
# View check output
sudo datadog-agent check <check_name> --check-rate