Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Datadog

Cloud-based monitoring and analytics platform for infrastructure, applications, logs, and traces with unified observability.

Datadog

Cloud-based monitoring and analytics platform for infrastructure, applications, logs, and traces with unified observability.

Overview

Datadog is a SaaS-based monitoring and analytics platform that provides full-stack observability across infrastructure, applications, logs, and traces. It uses an agent-based approach to collect metrics, logs, and traces from hosts and containers, sending data to Datadog's cloud platform for analysis, visualisation, and alerting. The platform integrates with hundreds of technologies and supports custom instrumentation.

Datadog PlatformCollection LayerData SourcesHosts/VMsContainersCloud ServicesApplicationsDatadog AgentDogStatsDAPI/SDKCloud IntegrationsMetrics StorageLog StorageAPM StorageQuery EngineDashboardsMonitors/AlertsAnalyticsDatadog PlatformCollection LayerData SourcesHosts/VMsContainersCloud ServicesApplicationsDatadog AgentDogStatsDAPI/SDKCloud IntegrationsMetrics StorageLog StorageAPM StorageQuery EngineDashboardsMonitors/AlertsAnalytics

Agent Installation and Configuration

The Datadog Agent is a lightweight daemon that collects metrics, logs, and traces from your hosts.

Installation

# Ubuntu/Debian, RedHat/CentOS/Amazon Linux (one script for all distros)
# The legacy s3.amazonaws.com/dd-agent install_script.sh is deprecated and
# defaults to Agent v6; use the versioned install.datadoghq.com script.
DD_API_KEY=<YOUR_API_KEY> DD_SITE="datadoghq.com" \
  bash -c "$(curl -L https://install.datadoghq.com/scripts/install_script_agent7.sh)"

# Docker
docker run -d --name datadog-agent \
  -e DD_API_KEY=<YOUR_API_KEY> \
  -e DD_SITE="datadoghq.com" \
  -v /var/run/docker.sock:/var/run/docker.sock:ro \
  -v /proc/:/host/proc/:ro \
  -v /sys/fs/cgroup/:/host/sys/fs/cgroup:ro \
  gcr.io/datadoghq/agent:7

# Kubernetes (Helm)
helm repo add datadog https://helm.datadoghq.com
helm install datadog-agent datadog/datadog \
  --set datadog.apiKey=<YOUR_API_KEY> \
  --set datadog.site=datadoghq.com

Main Configuration File

# /etc/datadog-agent/datadog.yaml
api_key: <YOUR_API_KEY>
site: datadoghq.com

# Hostname (optional, auto-detected by default)
hostname: web-server-01

# Tags for this host
tags:
  - env:production
  - role:webserver
  - team:platform

# Enable logs collection
logs_enabled: true

# Enable APM
apm_config:
  enabled: true
  env: production

# Enable process monitoring
process_config:
  enabled: true

# Network performance monitoring
network_config:
  enabled: true

# Enable live container monitoring
container_collect_all: true

Agent Commands

# Start/stop/restart agent
sudo systemctl start datadog-agent
sudo systemctl stop datadog-agent
sudo systemctl restart datadog-agent

# Check agent status
sudo datadog-agent status

# Check specific integration status
sudo datadog-agent status integrations

# Validate configuration
sudo datadog-agent configcheck

# Test connectivity
sudo datadog-agent diagnose

# Restart specific integration
sudo datadog-agent restart <integration_name>

Docker Agent Configuration

# docker-compose.yml
version: '3'
services:
  datadog-agent:
    image: gcr.io/datadoghq/agent:7
    environment:
      - DD_API_KEY=${DD_API_KEY}
      - DD_SITE=datadoghq.com
      - DD_LOGS_ENABLED=true
      - DD_APM_ENABLED=true
      - DD_APM_NON_LOCAL_TRAFFIC=true
      - DD_PROCESS_AGENT_ENABLED=true
      - DD_TAGS=env:production role:docker
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - /proc/:/host/proc/:ro
      - /sys/fs/cgroup:/host/sys/fs/cgroup:ro
      - /etc/datadog-agent/conf.d:/etc/datadog-agent/conf.d:ro
    ports:
      - "8126:8126/tcp"  # APM
      - "8125:8125/udp"  # DogStatsD

Metrics Collection

Datadog collects standard infrastructure metrics automatically and supports custom metrics through various methods.

Standard Metrics

# System metrics (collected automatically)
system.cpu.user           # CPU usage by user processes
system.cpu.system         # CPU usage by system processes
system.mem.used           # Memory used
system.mem.free           # Memory free
system.disk.used          # Disk space used
system.disk.free          # Disk space free
system.net.bytes_sent     # Network bytes sent
system.net.bytes_rcvd     # Network bytes received
system.load.1             # 1-minute load average
system.load.5             # 5-minute load average
system.load.15            # 15-minute load average

# Docker metrics (when Docker integration enabled)
docker.containers.running # Number of running containers
docker.containers.stopped # Number of stopped containers
docker.cpu.usage          # Container CPU usage
docker.mem.usage          # Container memory usage

Custom Metrics with DogStatsD

# Python - using datadog library
from datadog import initialize, statsd

options = {
    'statsd_host': '127.0.0.1',
    'statsd_port': 8125
}
initialize(**options)

# Counter - tracks cumulative count
statsd.increment('page.views', tags=['page:home', 'env:prod'])
statsd.increment('api.requests', value=5)

# Gauge - tracks current value
statsd.gauge('database.connections', 42, tags=['db:mysql'])
statsd.gauge('queue.size', 100)

# Histogram - statistical distribution
statsd.histogram('request.duration', 0.235, tags=['endpoint:/api/users'])

# Distribution - global percentiles
statsd.distribution('request.size', 1024, tags=['method:POST'])

# Set - count unique values
statsd.set('unique.visitors', 'user123', tags=['site:main'])

# Timing - convenience for histograms
statsd.timing('query.time', 350)  # milliseconds

# Context manager for timing
with statsd.timed('database.query', tags=['query:users']):
    # Your code here
    pass
// Node.js - using hot-shots library
const StatsD = require('hot-shots');
const dogstatsd = new StatsD({
    host: 'localhost',
    port: 8125,
    globalTags: ['env:prod']
});

// Counter
dogstatsd.increment('page.views', 1, ['page:home']);

// Gauge
dogstatsd.gauge('active.connections', 42);

// Histogram
dogstatsd.histogram('response.time', 235, ['endpoint:/api']);

// Distribution
dogstatsd.distribution('file.size', 2048);

// Timing
dogstatsd.timing('query.duration', 150);

// Set
dogstatsd.set('unique.users', 'user456');
// Go - using datadog-go library
import "github.com/DataDog/datadog-go/v5/statsd"

client, err := statsd.New("127.0.0.1:8125",
    statsd.WithNamespace("myapp."),
    statsd.WithTags([]string{"env:prod"}),
)
if err != nil {
    log.Fatal(err)
}
defer client.Close()

// Counter
client.Incr("page.views", []string{"page:home"}, 1)

// Gauge
client.Gauge("queue.size", 100, []string{"queue:jobs"}, 1)

// Histogram
client.Histogram("request.duration", 0.235, []string{"endpoint:/api"}, 1)

// Distribution
client.Distribution("request.size", 1024, []string{"method:POST"}, 1)

// Timing
client.Timing("query.time", time.Millisecond*350, []string{"db:postgres"}, 1)

Custom Metrics via Agent Checks

# /etc/datadog-agent/checks.d/custom_check.py
from datadog_checks.base import AgentCheck

class CustomCheck(AgentCheck):
    def check(self, instance):
        # Gauge metric
        self.gauge(
            'custom.metric.value',
            42,
            tags=['env:prod', 'service:api']
        )

        # Counter metric
        self.count(
            'custom.events.total',
            1,
            tags=['type:login']
        )

        # Rate metric (count per second)
        self.rate(
            'custom.requests.rate',
            10
        )

        # Service check
        self.service_check(
            'custom.service.status',
            AgentCheck.OK,
            message='Service is healthy',
            tags=['service:api']
        )
# /etc/datadog-agent/conf.d/custom_check.yaml
init_config:

instances:
  - min_collection_interval: 60  # Run every 60 seconds

Log Collection and Parsing

Datadog Agent can collect logs from various sources and parse them for analysis.

Enable Log Collection

# /etc/datadog-agent/datadog.yaml
logs_enabled: true

# Optional log processing
logs_config:
  # Process raw logs on agent before sending
  use_http: true
  use_compression: true
  compression_level: 6

  # Additional processing
  processing_rules:
    - type: exclude_at_match
      name: exclude_healthcheck
      pattern: /healthcheck

File-based Log Collection

# /etc/datadog-agent/conf.d/custom_logs.yaml
logs:
  - type: file
    path: /var/log/myapp/*.log
    service: myapp
    source: python
    sourcecategory: sourcecode
    tags:
      - env:production
      - team:backend

    # Multi-line log aggregation
    log_processing_rules:
      - type: multi_line
        name: log_start_pattern
        pattern: \d{4}-\d{2}-\d{2}  # Lines starting with date

  # JSON logs
  - type: file
    path: /var/log/myapp/json.log
    service: myapp
    source: python
    sourcecategory: sourcecode

  # Nginx access logs
  - type: file
    path: /var/log/nginx/access.log
    service: nginx
    source: nginx
    sourcecategory: http_web_access

Docker Log Collection

# docker-compose.yml labels
services:
  web:
    image: myapp:latest
    labels:
      com.datadoghq.ad.logs: '[{"source": "python", "service": "myapp"}]'
      com.datadoghq.ad.tags: '["env:prod", "version:1.2.3"]'
# Kubernetes pod annotations
apiVersion: v1
kind: Pod
metadata:
  name: myapp
  annotations:
    ad.datadoghq.com/myapp.logs: |
      [
        {
          "source": "python",
          "service": "myapp",
          "tags": ["env:prod"]
        }
      ]
spec:
  containers:
    - name: myapp
      image: myapp:latest

Log Processing Rules

# /etc/datadog-agent/conf.d/custom_logs.yaml
logs:
  - type: file
    path: /var/log/myapp/app.log
    service: myapp
    source: python

    log_processing_rules:
      # Exclude logs matching pattern
      - type: exclude_at_match
        name: exclude_debug
        pattern: DEBUG

      # Include only logs matching pattern
      - type: include_at_match
        name: include_errors
        pattern: ERROR|CRITICAL

      # Mask sensitive data
      - type: mask_sequences
        name: mask_credit_cards
        pattern: \d{4}-\d{4}-\d{4}-\d{4}
        replace_placeholder: "[MASKED_CC]"

      # Multi-line aggregation
      - type: multi_line
        name: aggregate_stack_traces
        pattern: ^\s+at

Log Pipeline and Parsing

// Grok parser example (configured in Datadog UI)
// Pattern for parsing custom log format
rule %{date("yyyy-MM-dd HH:mm:ss"):timestamp} \[%{word:log_level}\] %{data:logger} - %{data:message}

// Example log:
// 2025-12-03 10:30:45 [ERROR] app.handlers - Database connection failed

// Parsed attributes:
// timestamp: 2025-12-03 10:30:45
// log_level: ERROR
// logger: app.handlers
// message: Database connection failed

Common Log Patterns

# Application logs
logs:
  - type: file
    path: /var/log/app/application.log
    service: myapp
    source: java
    log_processing_rules:
      - type: multi_line
        name: java_stack_trace
        pattern: ^\s+(at|\.{3})\s+

# Syslog
logs:
  - type: tcp
    port: 10514
    service: syslog
    source: syslog

# JSON logs (auto-parsed)
logs:
  - type: file
    path: /var/log/app/json.log
    service: myapp
    source: nodejs
    # JSON logs are automatically parsed

APM and Distributed Tracing

Application Performance Monitoring tracks requests across services with distributed tracing.

DatadogDatadogAgentDatabaseAPIServiceWebServiceClientDatadogDatadogAgentDatabaseAPIServiceWebServiceClientHTTP RequestInternal API CallQueryResultResponseHTTP ResponseTrace SpanTrace SpanTracesDatadogDatadogAgentDatabaseAPIServiceWebServiceClientDatadogDatadogAgentDatabaseAPIServiceWebServiceClientHTTP RequestInternal API CallQueryResultResponseHTTP ResponseTrace SpanTrace SpanTraces

Enable APM Agent

# /etc/datadog-agent/datadog.yaml
apm_config:
  enabled: true
  env: production

  # Allow traces from other containers
  apm_non_local_traffic: true

  # Ingestion sampling (current mechanism; the legacy App Analytics
  # `analyzed_rate_by_service` is superseded).
  # Head-based per-service/resource sampling rules in the Agent:
  trace_sampling_rules:
    - service: my-service
      sample_rate: 1.0   # Keep 100% of this service's traces
  # Per-host throughput cap for the Agent's automatic sampler (env: DD_APM_MAX_TPS):
  max_traces_per_second: 50
  # Tracing-library-side rules use DD_TRACE_SAMPLING_RULES instead.
  # For volume management, also see Datadog ingestion controls + retention filters.

  # Resource filtering
  filter_tags:
    require: ["env:production"]

  # Obfuscation
  obfuscation:
    elasticsearch:
      enabled: true
    mongodb:
      enabled: true
    http:
      remove_query_string: true
      remove_paths_with_digits: true
    redis:
      enabled: true

Python APM Instrumentation

# Automatic instrumentation
from ddtrace import patch_all
patch_all()

# Application code
from flask import Flask
app = Flask(__name__)

@app.route('/')
def hello():
    return 'Hello World!'

if __name__ == '__main__':
    app.run()

# Run with ddtrace
# DD_SERVICE=myapp DD_ENV=prod DD_VERSION=1.0 \
# ddtrace-run python app.py
# Manual instrumentation
from ddtrace import tracer

@tracer.wrap(service='myapp', resource='process_data')
def process_data(data):
    with tracer.trace('database.query', service='postgres') as span:
        span.set_tag('query.type', 'SELECT')
        span.set_tag('rows.count', 100)
        result = db.query(data)
    return result

# Custom span
with tracer.trace('custom.operation') as span:
    span.set_tag('user.id', user_id)
    span.set_metric('items.processed', 42)
    # Your code here

Node.js APM Instrumentation

// tracer.js - initialize first
const tracer = require('dd-trace').init({
  service: 'myapp',
  env: 'production',
  version: '1.0.0',
  analytics: true,
  logInjection: true
});

module.exports = tracer;
// app.js
require('./tracer'); // Must be first

const express = require('express');
const app = express();

// Automatic instrumentation for supported libraries
app.get('/', (req, res) => {
  res.send('Hello World!');
});

// Manual instrumentation
const tracer = require('dd-trace');

function processOrder(order) {
  const span = tracer.startSpan('process.order', {
    tags: {
      'order.id': order.id,
      'order.amount': order.amount
    }
  });

  try {
    // Your code here
    span.setTag('result', 'success');
  } catch (err) {
    span.setTag('error', true);
    span.setTag('error.msg', err.message);
    throw err;
  } finally {
    span.finish();
  }
}

Go APM Instrumentation

package main

import (
    "net/http"
    httptrace "gopkg.in/DataDog/dd-trace-go.v1/contrib/net/http"
    "gopkg.in/DataDog/dd-trace-go.v1/ddtrace/tracer"
)

func main() {
    // Initialize tracer
    tracer.Start(
        tracer.WithService("myapp"),
        tracer.WithEnv("production"),
        tracer.WithServiceVersion("1.0.0"),
    )
    defer tracer.Stop()

    // Automatic HTTP instrumentation
    mux := httptrace.NewServeMux()
    mux.HandleFunc("/", handler)
    http.ListenAndServe(":8080", mux)
}

func handler(w http.ResponseWriter, r *http.Request) {
    // Manual span creation
    span, ctx := tracer.StartSpanFromContext(r.Context(), "custom.operation")
    defer span.Finish()

    span.SetTag("user.id", "123")

    // Your code here
    processData(ctx)
}

func processData(ctx context.Context) {
    span, _ := tracer.StartSpanFromContext(ctx, "process.data")
    defer span.Finish()

    // Your code here
}

Java APM Instrumentation

# Download Java agent
wget -O dd-java-agent.jar \
  https://dtdg.co/latest-java-tracer

# Run application with agent
java -javaagent:dd-java-agent.jar \
  -Ddd.service=myapp \
  -Ddd.env=production \
  -Ddd.version=1.0.0 \
  -Ddd.trace.analytics.enabled=true \
  -jar myapp.jar

Service Map Configuration

# Tag spans for service mapping
# Tags are automatically detected from:
# - service: service name
# - env: environment
# - version: application version
# - resource: specific endpoint or operation

# Example span tags
span.service = "web-frontend"
span.env = "production"
span.version = "2.1.0"
span.resource = "GET /api/users"
span.type = "web"

Dashboard Creation

Datadog dashboards provide customisable visualisations for metrics, logs, and traces.

Dashboard Types

Type Description Use Case
Timeboard Fixed time synchronisation across all widgets System monitoring, correlation analysis
Screenboard Free-form layout with independent time ranges Executive dashboards, status boards

Common Widget Types

// Timeseries widget
{
  "definition": {
    "type": "timeseries",
    "requests": [
      {
        "q": "avg:system.cpu.user{*}",
        "display_type": "line",
        "style": {
          "palette": "dog_classic",
          "line_type": "solid",
          "line_width": "normal"
        }
      }
    ],
    "title": "CPU Usage",
    "show_legend": true,
    "legend_size": "0"
  }
}

// Query value widget (single metric)
{
  "definition": {
    "type": "query_value",
    "requests": [
      {
        "q": "avg:system.mem.used{*}",
        "aggregator": "avg"
      }
    ],
    "title": "Memory Used",
    "precision": 2,
    "autoscale": true
  }
}

// Heatmap widget
{
  "definition": {
    "type": "heatmap",
    "requests": [
      {
        "q": "avg:trace.http.request.duration{*} by {resource_name}"
      }
    ],
    "title": "Request Duration Heatmap"
  }
}

// Top list widget
{
  "definition": {
    "type": "toplist",
    "requests": [
      {
        "q": "top(avg:docker.cpu.usage{*} by {container_name}, 10, 'mean', 'desc')"
      }
    ],
    "title": "Top 10 Containers by CPU"
  }
}

Dashboard via API

from datadog import initialize, api

initialize(api_key='<YOUR_API_KEY>', app_key='<YOUR_APP_KEY>')

# Create dashboard
dashboard = {
    'title': 'System Overview',
    'description': 'System metrics dashboard',
    'layout_type': 'ordered',
    'widgets': [
        {
            'definition': {
                'type': 'timeseries',
                'requests': [
                    {
                        'q': 'avg:system.cpu.user{*}',
                        'display_type': 'line'
                    }
                ],
                'title': 'CPU Usage'
            }
        },
        {
            'definition': {
                'type': 'query_value',
                'requests': [
                    {
                        'q': 'avg:system.mem.used{*}',
                        'aggregator': 'last'
                    }
                ],
                'title': 'Memory Used',
                'autoscale': True
            }
        }
    ]
}

result = api.Dashboard.create(**dashboard)
print(f"Dashboard created: {result['id']}")

Template Variables

// Dashboard with template variables
{
  "title": "Infrastructure Dashboard",
  "template_variables": [
    {
      "name": "env",
      "default": "production",
      "prefix": "env"
    },
    {
      "name": "host",
      "default": "*",
      "prefix": "host"
    },
    {
      "name": "service",
      "default": "*",
      "prefix": "service"
    }
  ],
  "widgets": [
    {
      "definition": {
        "type": "timeseries",
        "requests": [
          {
            "q": "avg:system.cpu.user{$env,$host,$service}"
          }
        ]
      }
    }
  ]
}

Monitor and Alert Configuration

Monitors detect conditions and trigger alerts based on metric thresholds and anomalies.

Threshold ExceededWithin BoundsMetric NormalMetric DataMonitor EvaluationAlert StateOK StateNotificationChannelsEmailSlackPagerDutyWebhookNo ActionAlert RecoveryThreshold ExceededWithin BoundsMetric NormalMetric DataMonitor EvaluationAlert StateOK StateNotificationChannelsEmailSlackPagerDutyWebhookNo ActionAlert Recovery

Monitor Types

Type Description Use Case
Metric Alert on metric threshold CPU > 80%, memory usage
APM Alert on trace metrics Error rate, latency p99
Integration Cloud service monitoring AWS RDS, ELB status
Process Monitor process availability Nginx running, pod count
Network Network performance TCP connections, bandwidth
Log Alert on log patterns Error log frequency
Event Alert on custom events Deployment events
Composite Combine multiple monitors Complex conditions
Anomaly ML-based anomaly detection Unusual traffic patterns
Outlier Detect outlier hosts One host behaving differently
Forecast Predict future values Disk will fill in 2 days

Metric Monitor

// Simple threshold monitor
{
  "name": "High CPU Usage",
  "type": "metric alert",
  "query": "avg(last_5m):avg:system.cpu.user{env:production} > 80",
  "message": "CPU usage is above 80% on {{host.name}} @slack-alerts",
  "tags": ["team:infrastructure", "priority:high"],
  "options": {
    "notify_audit": false,
    "locked": false,
    "timeout_h": 0,
    "include_tags": true,
    "no_data_timeframe": 10,
    "require_full_window": false,
    "new_host_delay": 300,
    "notify_no_data": true,
    "renotify_interval": 0,
    "escalation_message": "CPU still high after 30 minutes",
    "thresholds": {
      "critical": 80,
      "warning": 70,
      "critical_recovery": 75,
      "warning_recovery": 65
    }
  }
}

Multi-Alert Monitor

// Alert per service
{
  "name": "High Error Rate per Service",
  "type": "metric alert",
  "query": "avg(last_10m):sum:trace.http.request.errors{env:production} by {service}.as_rate() > 5",
  "message": "Error rate for {{service.name}} is above 5% @pagerduty",
  "options": {
    "thresholds": {
      "critical": 5,
      "warning": 2
    },
    "notify_no_data": false,
    "no_data_timeframe": 20
  }
}

Anomaly Detection Monitor

// Anomaly detection for traffic patterns
{
  "name": "Anomalous Request Rate",
  "type": "query alert",
  "query": "avg(last_4h):anomalies(avg:trace.http.request.hits{env:production}.as_rate(), 'basic', 2, direction='both', alert_window='last_15m', interval=60, count_default_zero='true') >= 1",
  "message": "Unusual request rate detected @slack-alerts",
  "options": {
    "thresholds": {
      "critical": 1,
      "critical_recovery": 0
    },
    "threshold_windows": {
      "trigger_window": "last_15m",
      "recovery_window": "last_15m"
    }
  }
}

Log Monitor

// Alert on error log frequency
{
  "name": "High Error Log Rate",
  "type": "log alert",
  "query": "logs(\"status:error service:myapp\").index(\"main\").rollup(\"count\").last(\"5m\") > 100",
  "message": "More than 100 errors in last 5 minutes @slack-alerts\n{{#is_alert}}\nError details: {{log.message}}\n{{/is_alert}}",
  "options": {
    "thresholds": {
      "critical": 100,
      "warning": 50
    },
    "enable_logs_sample": true
  }
}

Composite Monitor

# Create composite monitor via API
from datadog import initialize, api

initialize(api_key='<API_KEY>', app_key='<APP_KEY>')

# Composite monitor combining CPU and memory
monitor = {
    'name': 'High Resource Usage',
    'type': 'composite',
    'query': '(monitor_id_1 && monitor_id_2) || monitor_id_3',
    'message': 'Critical resource usage detected @pagerduty',
    'options': {
        'notify_no_data': False
    }
}

api.Monitor.create(**monitor)

Monitor Templates

// Monitor with template variables
{
  "name": "{{service.name}} - High Latency",
  "type": "metric alert",
  "query": "avg(last_5m):avg:trace.http.request.duration.by.service.99p{service:{{service.name}},env:production} > 1",
  "message": "P99 latency for {{service.name}} is above 1s\n\nHost: {{host.name}}\nEnv: {{env.name}}\n\n@slack-{{service.name}}",
  "tags": ["service:{{service.name}}", "auto-generated"]
}

Notification Channels

# Alert message format with @ mentions
@slack-alerts           # Slack channel
@pagerduty-critical     # PagerDuty service
@webhook-custom         # Custom webhook
@user@example.com       # Email address
@all                    # All team members

# Conditional notifications
{{#is_alert}}
Alert triggered!
{{/is_alert}}

{{#is_warning}}
Warning level reached
{{/is_warning}}

{{#is_recovery}}
Issue resolved
{{/is_recovery}}

{{#is_no_data}}
No data received
{{/is_no_data}}

Monitor via API

from datadog import initialize, api

initialize(api_key='<API_KEY>', app_key='<APP_KEY>')

# Create monitor
monitor = {
    'name': 'Disk Space Low',
    'type': 'metric alert',
    'query': 'avg(last_5m):avg:system.disk.free{*} by {host} < 10000000000',
    'message': '''Disk space below 10GB on {{host.name}}

Current value: {{value}}
Threshold: {{threshold}}

@slack-infrastructure''',
    'tags': ['auto:true', 'team:platform'],
    'options': {
        'thresholds': {
            'critical': 10000000000,
            'warning': 20000000000
        },
        'notify_no_data': True,
        'no_data_timeframe': 20,
        'require_full_window': False,
        'notify_audit': False,
        'include_tags': True
    }
}

result = api.Monitor.create(**monitor)
print(f"Monitor created: {result['id']}")

# Update monitor
api.Monitor.update(
    monitor_id,
    query='avg(last_5m):avg:system.disk.free{*} by {host} < 5000000000'
)

# Delete monitor
api.Monitor.delete(monitor_id)

# Mute monitor
api.Monitor.mute(monitor_id, end=1735948800)  # Unix timestamp

# Unmute monitor
api.Monitor.unmute(monitor_id)

Common Integrations

Datadog integrates with hundreds of services and platforms for automatic metric collection.

AWS Integration

# Configure via Datadog UI or API
# Requires IAM role with permissions:
# - cloudwatch:GetMetricStatistics
# - cloudwatch:ListMetrics
# - ec2:DescribeInstances
# - ec2:DescribeTags
# - rds:DescribeDBInstances
# - s3:ListBucket

# Automatically collects:
# - EC2 instance metrics
# - RDS database metrics
# - ELB load balancer metrics
# - Lambda function metrics
# - S3 bucket metrics
# - CloudWatch custom metrics

Kubernetes Integration

# values.yaml for Helm chart
datadog:
  apiKey: <YOUR_API_KEY>
  site: datadoghq.com

  # Cluster monitoring
  clusterName: production-cluster

  # Collect Kubernetes events
  collectEvents: true

  # Leader election for cluster checks
  leaderElection: true

  # Kubernetes state metrics
  kubeStateMetricsEnabled: true

  # APM
  apm:
    enabled: true
    port: 8126

  # Logs
  logs:
    enabled: true
    containerCollectAll: true

  # Process monitoring
  processAgent:
    enabled: true

  # Network performance monitoring
  networkMonitoring:
    enabled: true

  # Cluster checks
  clusterChecks:
    enabled: true

# Kubernetes metrics collected automatically:
# - kubernetes.cpu.usage
# - kubernetes.memory.usage
# - kubernetes.network.rx_bytes
# - kubernetes.network.tx_bytes
# - kubernetes_state.deployment.replicas_desired
# - kubernetes_state.pod.ready

PostgreSQL Integration

# /etc/datadog-agent/conf.d/postgres.d/conf.yaml
init_config:

instances:
  - host: localhost
    port: 5432
    username: datadog
    password: '<PASSWORD>'
    dbname: postgres

    # Collect additional metrics
    collect_function_metrics: true
    collect_count_metrics: true
    collect_activity_metrics: true
    collect_database_size_metrics: true
    collect_default_database: true

    # Custom queries
    custom_queries:
      - metric_prefix: postgresql
        query: SELECT count(*) as connections FROM pg_stat_activity;
        columns:
          - name: connections
            type: gauge
        tags:
          - custom_query:active_connections

    tags:
      - env:production
      - db:main
-- Create Datadog user in PostgreSQL
CREATE USER datadog WITH PASSWORD '<PASSWORD>';
GRANT pg_monitor TO datadog;
GRANT SELECT ON pg_stat_database TO datadog;

Redis Integration

# /etc/datadog-agent/conf.d/redisdb.d/conf.yaml
init_config:

instances:
  - host: localhost
    port: 6379
    password: '<PASSWORD>'

    # Collect slow commands
    command_stats: true

    # Warn on slow operations
    slowlog-max-len: 128

    tags:
      - env:production
      - cache:main

NGINX Integration

# /etc/datadog-agent/conf.d/nginx.d/conf.yaml
init_config:

instances:
  - nginx_status_url: http://localhost/nginx_status/

    # Additional configurations
    tags:
      - env:production
      - instance:web-1
# Enable NGINX stub_status module
# /etc/nginx/sites-available/default
server {
    listen 80;

    location /nginx_status {
        stub_status on;
        access_log off;
        allow 127.0.0.1;
        deny all;
    }
}

MySQL Integration

# /etc/datadog-agent/conf.d/mysql.d/conf.yaml
init_config:

instances:
  - host: localhost
    port: 3306
    user: datadog
    pass: '<PASSWORD>'

    # Replication metrics
    replication: true

    # InnoDB metrics
    extra_innodb_metrics: true

    # Performance schema
    extra_performance_metrics: true

    # Schema metrics
    schema_size_metrics: true

    tags:
      - env:production
      - db:primary
-- Create Datadog user in MySQL
CREATE USER 'datadog'@'localhost' IDENTIFIED BY '<PASSWORD>';
GRANT REPLICATION CLIENT ON *.* TO 'datadog'@'localhost';
GRANT PROCESS ON *.* TO 'datadog'@'localhost';
GRANT SELECT ON performance_schema.* TO 'datadog'@'localhost';

Docker Integration

# Automatically enabled when Docker socket is mounted
# Collects container metrics:
# - docker.cpu.usage
# - docker.mem.usage
# - docker.io.read_bytes
# - docker.io.write_bytes
# - docker.net.bytes_sent
# - docker.net.bytes_rcvd

# docker-compose.yml
services:
  datadog-agent:
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
    environment:
      - DD_CONTAINER_EXCLUDE: "name:datadog-agent"  # Exclude self
      - DD_CONTAINER_INCLUDE: "image:myapp.*"       # Include specific images

Query Language Basics

Datadog uses a query language for metrics, logs, and traces.

Metric Query Syntax

# Basic structure
function(aggregation:metric_name{tag_filters}[.rollup(method, time)])

# Examples
avg:system.cpu.user{host:web-01}
sum:http.requests{service:api,env:prod}.as_count()
max:system.mem.used{*} by {host}
p95:trace.http.request.duration{service:web}

# Aggregation functions
avg     # Average
sum     # Sum
min     # Minimum
max     # Maximum
count   # Count

# Time aggregation (rollup)
.as_count()      # Sum over rollup period
.as_rate()       # Per-second rate
.rollup(avg, 60) # Average over 60 seconds

Metric Functions

# Arithmetic
avg:system.cpu.user{*} + avg:system.cpu.system{*}
avg:system.mem.used{*} / avg:system.mem.total{*} * 100

# Rate of change
per_second(sum:http.requests{*}.as_count())
per_minute(sum:errors.total{*})

# Derivative (change between consecutive points)
derivative(avg:system.disk.used{*})

# Moving average
ewma_10(avg:system.load.1{*})     # Exponentially weighted moving average
autosmooth(avg:latency{*})         # Automatic smoothing

# Anomaly detection
anomalies(avg:requests.count{*}, 'basic', 2)
anomalies(avg:cpu.usage{*}, 'agile', 3, direction='above')

# Forecasting
forecast(avg:disk.used{*}, 'linear', 2)  # Forecast 2 periods ahead

# Rollup functions
.rollup(avg, 300)    # 5-minute average
.rollup(sum, 3600)   # 1-hour sum
.rollup(max, 60)     # 1-minute maximum

# Top/bottom
top(avg:cpu.usage{*} by {host}, 5, 'mean', 'desc')
bottom(avg:latency{*} by {endpoint}, 3, 'last', 'asc')

Tag Filtering

# Exact match
metric{tag:value}
metric{env:production}

# Multiple tags (AND)
metric{env:production,service:api}

# Wildcard
metric{host:web-*}
metric{service:*api*}

# OR condition
metric{env:production OR env:staging}

# NOT condition
metric{env:production,host:!web-01}

# Group by tags
avg:metric{*} by {host}
sum:metric{*} by {service,env}

Log Query Syntax

# Basic search
status:error
service:myapp
message:"connection timeout"

# Boolean operators
status:error AND service:myapp
status:(error OR warning)
status:error AND NOT service:healthcheck

# Wildcards
service:web-*
message:*timeout*

# Numeric ranges
http.status_code:[400 TO 499]
duration:>1000
response_time:[100 TO 500]

# Facet search
@http.method:POST
@user.id:12345
@error.kind:DatabaseError

# Nested fields (JSON)
@attributes.user.email:"user@example.com"

# Existence check
@http.status_code:*  # Has status code field
-@http.status_code:* # Missing status code field

Log Aggregations

# Count
count() by @service

# Measure
avg(@duration) by @http.method
max(@response_time) by @endpoint
p95(@latency) by @region

# Multiple dimensions
count() by @service,@env
avg(@duration) by @service,@http.status_code

APM Query Syntax

# Trace search
service:web-frontend
resource:"GET /api/users"
env:production

# Error traces
error:true
@http.status_code:[500 TO 599]

# Duration filtering
duration:>1s
@duration:[100ms TO 500ms]

# Resource patterns
resource:"GET /api/*"
operation:http.request

# Span tags
@http.method:POST
@db.type:postgres
@user.id:12345

# Analytics
avg(@duration) by service
p99(@duration) by resource
count() by @http.status_code

Time Range Functions

# Relative time ranges
last_5m      # Last 5 minutes
last_1h      # Last 1 hour
last_4h      # Last 4 hours
last_1d      # Last 1 day
last_1w      # Last 1 week

# Compare to previous period
avg:metric{*}, avg:metric{*}.shift(-1w)  # Compare to last week
avg:metric{*}, avg:metric{*}.shift(-1d)  # Compare to yesterday

# Time shift
.shift(-1h)   # Shift back 1 hour
.shift(1d)    # Shift forward 1 day

Quick Reference

Essential Commands

# Agent management
sudo systemctl start datadog-agent
sudo systemctl stop datadog-agent
sudo systemctl restart datadog-agent
sudo datadog-agent status
sudo datadog-agent configcheck

# Testing
sudo datadog-agent check <integration_name>
sudo datadog-agent diagnose
sudo -u dd-agent datadog-agent check <check_name>

# Logs
sudo tail -f /var/log/datadog/agent.log
sudo journalctl -u datadog-agent -f

Common Metrics

# System
system.cpu.user              # CPU user time
system.cpu.system            # CPU system time
system.mem.used              # Memory used
system.disk.used             # Disk used
system.net.bytes_sent        # Network sent
system.load.1                # 1-min load average

# Docker
docker.cpu.usage             # Container CPU
docker.mem.usage             # Container memory
docker.containers.running    # Running containers

# APM
trace.http.request.hits      # Request count
trace.http.request.errors    # Error count
trace.http.request.duration  # Request duration

Key Configuration Files

# Main configuration
/etc/datadog-agent/datadog.yaml

# Integration configs
/etc/datadog-agent/conf.d/<integration>.d/conf.yaml

# Custom checks
/etc/datadog-agent/checks.d/
/etc/datadog-agent/conf.d/

# Logs
/var/log/datadog/agent.log
/var/log/datadog/trace-agent.log

API Example

from datadog import initialize, api

initialize(api_key='<API_KEY>', app_key='<APP_KEY>')

# Post metric
api.Metric.send(
    metric='custom.metric',
    points=[(time.time(), 42)],
    tags=['env:prod']
)

# Query metrics
api.Metric.query(
    start=int(time.time()) - 3600,
    end=int(time.time()),
    query='avg:system.cpu.user{*}'
)

# Create monitor
api.Monitor.create(
    type='metric alert',
    query='avg(last_5m):avg:system.cpu.user{*} > 80',
    name='High CPU',
    message='@slack-alerts'
)

Common Issues and Solutions

Issue Cause Solution
Agent not reporting API key incorrect Verify API key in datadog.yaml, check datadog-agent status
High agent CPU usage Too many integrations/checks Reduce check frequency, disable unused integrations
Missing metrics Integration not configured Check integration config in /etc/datadog-agent/conf.d/
No logs appearing Logs not enabled Set logs_enabled: true in datadog.yaml
APM traces missing Agent APM disabled Enable with apm_config.enabled: true
Service not traced Library not instrumented Install and configure APM library for language
High data usage Too many custom metrics Use metric aggregation, reduce cardinality, filter tags
Duplicate metrics Multiple agents reporting Check hostname configuration, ensure unique hostnames
Monitor flapping Threshold too sensitive Adjust threshold or increase evaluation window
No data in dashboards Wrong time range selected Verify time range and template variable selections
Container metrics missing Docker socket not mounted Mount /var/run/docker.sock in agent container
Kubernetes metrics missing Agent not deployed as DaemonSet Deploy agent on every node, enable cluster checks
Slow dashboard loading Too many queries Reduce number of widgets, increase time granularity
API rate limit exceeded Too many API calls Implement rate limiting, batch requests
Logs not parsed Missing pipeline Create log pipeline with parsers in Datadog UI
High latency in traces Sampling rate too high Reduce sampling rate in APM config

Debugging Steps

# Check agent status
sudo datadog-agent status

# Test specific integration
sudo datadog-agent check <integration_name> -l debug

# Check connectivity
sudo datadog-agent diagnose

# View recent logs
sudo tail -f /var/log/datadog/agent.log

# Check configuration validity
sudo datadog-agent configcheck

# List running checks
sudo datadog-agent check list

# Flare (send diagnostics to support)
sudo datadog-agent flare <case-id>

Performance Optimisation

# /etc/datadog-agent/datadog.yaml
# Reduce collection frequency
check_runners: 4

# Increase batch size for metrics
aggregator_buffer_size: 100

# Limit container collection
container_exclude: ["name:.*-test", "image:.*debug.*"]
container_include: ["name:prod-.*"]

# Reduce log collection
logs_config:
  # Increase batch size
  batch_wait: 5

  # Compress logs
  use_compression: true
  compression_level: 6

Custom Check Troubleshooting

# Test custom check
sudo -u dd-agent datadog-agent check <check_name> -l debug

# Check Python syntax
python3 -m py_compile /etc/datadog-agent/checks.d/<check_name>.py

# View check output
sudo datadog-agent check <check_name> --check-rate