Prometheus
Time-series database and monitoring system for collecting, storing, and querying metrics with powerful alerting capabilities.
Prometheus
Time-series database and monitoring system for collecting, storing, and querying metrics with powerful alerting capabilities.
Overview
Prometheus is an open-source systems monitoring and alerting toolkit that collects metrics from configured targets at specified intervals, evaluates rule expressions, displays results, and triggers alerts when specified conditions are met. It uses a pull-based model where Prometheus scrapes HTTP endpoints that expose metrics in a specific format.
graph TB
subgraph "Prometheus Architecture"
A[Service Discovery] --> B[Prometheus Server]
C[Exporters/Apps] -->|/metrics| B
B --> D[(TSDB Storage)]
B --> E[PromQL Engine]
E --> F[Grafana/API]
B --> G[Alertmanager]
G --> H[Email/Slack/PagerDuty]
end
subgraph "Targets"
I[Node Exporter]
J[Application]
K[Database Exporter]
end
I --> C
J --> C
K --> C
Metric Types
Prometheus supports four core metric types, each suited for different use cases.
Key Concepts
| Metric Type | Description | Use Case |
|---|---|---|
| Counter | Cumulative value that only increases or resets to zero | Request counts, errors, completed tasks |
| Gauge | Value that can go up or down | Temperature, memory usage, queue size |
| Histogram | Samples observations and counts them in buckets | Request latencies, response sizes |
| Summary | Similar to histogram but calculates quantiles client-side | Legacy use, when quantiles are pre-determined |
Common Patterns
# Counter - tracks cumulative values
http_requests_total{method="GET", status="200"} 1234
# Gauge - tracks current values
node_memory_available_bytes 4294967296
# Histogram - creates multiple time series
# _bucket: cumulative counts for each bucket
# _count: total number of observations
# _sum: sum of all observed values
http_request_duration_seconds_bucket{le="0.1"} 500
http_request_duration_seconds_bucket{le="0.5"} 800
http_request_duration_seconds_bucket{le="1.0"} 900
http_request_duration_seconds_bucket{le="+Inf"} 1000
http_request_duration_seconds_count 1000
http_request_duration_seconds_sum 450.5
# Summary - pre-calculated quantiles
http_request_duration_seconds{quantile="0.5"} 0.05
http_request_duration_seconds{quantile="0.9"} 0.1
http_request_duration_seconds{quantile="0.99"} 0.5
Examples
# Python client library examples
from prometheus_client import Counter, Gauge, Histogram, Summary
# Counter - increment only
requests_total = Counter(
'http_requests_total',
'Total HTTP requests',
['method', 'status']
)
requests_total.labels(method='GET', status='200').inc()
# Gauge - set to any value
temperature = Gauge('temperature_celsius', 'Current temperature')
temperature.set(21.5)
temperature.inc() # Increment by 1
temperature.dec() # Decrement by 1
# Histogram - observe values into buckets
request_latency = Histogram(
'request_latency_seconds',
'Request latency',
buckets=[0.1, 0.25, 0.5, 1.0, 2.5, 5.0]
)
request_latency.observe(0.35)
# Summary - calculate quantiles
response_size = Summary(
'response_size_bytes',
'Response size in bytes'
)
response_size.observe(512)
PromQL Basics
PromQL (Prometheus Query Language) is a powerful functional query language for selecting and aggregating time series data.
Key Concepts
flowchart LR
A[Instant Vector] -->|"rate()"| B[Instant Vector]
A -->|"[5m]"| C[Range Vector]
C -->|"avg_over_time()"| D[Instant Vector]
D -->|"sum by (label)"| E[Aggregated Vector]
Vector Types:
- Instant Vector: Single sample for each time series at a given timestamp
- Range Vector: Set of samples over a time range for each time series
- Scalar: Simple numeric floating point value
Selectors
# Exact match
http_requests_total{job="api-server"}
# Regex match
http_requests_total{method=~"GET|POST"}
# Negative match
http_requests_total{status!="500"}
# Negative regex match
http_requests_total{method!~"DELETE|PATCH"}
# Multiple conditions
http_requests_total{job="api-server", method="GET", status=~"2.."}
Operators
# Arithmetic operators
node_memory_total_bytes - node_memory_available_bytes
node_disk_written_bytes_total / 1024 / 1024 # Convert to MB
# Comparison operators (filter results)
http_requests_total > 100
node_filesystem_avail_bytes < 1e9
# Comparison returning boolean (keep value)
http_requests_total > bool 100
# Logical operators
up{job="api"} and on(instance) node_cpu_seconds_total
vector(1) or vector(2)
http_requests_total unless http_requests_total{status="200"}
Functions
# Rate functions (use with counters)
rate(http_requests_total[5m]) # Per-second rate over 5 minutes
irate(http_requests_total[5m]) # Instant rate (last two points)
increase(http_requests_total[1h]) # Total increase over 1 hour
# Aggregation over time (use with gauges)
avg_over_time(node_cpu_seconds_total[5m])
max_over_time(node_memory_usage_bytes[1h])
min_over_time(temperature_celsius[24h])
# Aggregation operators
sum(rate(http_requests_total[5m])) by (method)
avg(node_cpu_seconds_total) by (instance)
count(up) by (job)
topk(5, rate(http_requests_total[5m]))
bottomk(3, node_memory_available_bytes)
quantile(0.95, http_request_duration_seconds)
# Label manipulation
label_replace(up, "host", "$1", "instance", "(.*):.*")
label_join(up, "full_path", "/", "job", "instance")
# Math functions
abs(rate(errors_total[5m]))
ceil(memory_usage_bytes / 1024 / 1024)
floor(temperature_celsius)
round(latency_seconds, 0.001)
# Time functions
time() # Current Unix timestamp
timestamp(up) # Timestamp of each sample
day_of_week() # 0-6 (Sunday = 0)
hour() # 0-23
Examples
# Request rate per second by endpoint
sum(rate(http_requests_total[5m])) by (endpoint)
# Error rate percentage
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) * 100
# 95th percentile latency from histogram
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# Memory usage percentage
(node_memory_total_bytes - node_memory_available_bytes)
/ node_memory_total_bytes * 100
# CPU usage percentage (excluding idle)
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
Configuration
Prometheus configuration is defined in YAML format, controlling scrape targets, rules, and alerting.
Key Concepts
The main configuration file (prometheus.yml) contains:
- Global settings: Scrape interval, evaluation interval, external labels
- Scrape configs: Target definitions and how to scrape them
- Rule files: References to alerting and recording rule files
- Alerting config: Alertmanager connection settings
Common Patterns
# prometheus.yml
global:
scrape_interval: 15s # Default scrape interval
evaluation_interval: 15s # Rule evaluation frequency
scrape_timeout: 10s # Per-target scrape timeout
external_labels:
cluster: 'production'
region: 'eu-west-1'
# Alertmanager configuration
alerting:
alertmanagers:
- static_configs:
- targets:
- alertmanager:9093
# Rule files to load
rule_files:
- '/etc/prometheus/rules/*.yml'
- '/etc/prometheus/alerts/*.yml'
# Scrape configurations
scrape_configs:
# Prometheus self-monitoring
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
# Node exporters
- job_name: 'node'
static_configs:
- targets:
- 'node1:9100'
- 'node2:9100'
- 'node3:9100'
labels:
env: 'production'
# Application with custom path and scheme
- job_name: 'api-service'
scheme: https
metrics_path: '/internal/metrics'
basic_auth:
username: 'prometheus'
password_file: '/etc/prometheus/password'
tls_config:
ca_file: '/etc/prometheus/ca.crt'
insecure_skip_verify: false
static_configs:
- targets: ['api.example.com:443']
# Relabelling example
- job_name: 'kubernetes-pods'
relabel_configs:
# Keep only pods with annotation
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
# Use custom port from annotation
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
action: replace
target_label: __address__
regex: (.+)
replacement: $1
Examples
# Scrape config with metric relabelling
scrape_configs:
- job_name: 'filtered-metrics'
static_configs:
- targets: ['app:8080']
metric_relabel_configs:
# Drop specific metrics
- source_labels: [__name__]
regex: 'go_gc_.*'
action: drop
# Rename metric
- source_labels: [__name__]
regex: 'http_requests_total'
target_label: __name__
replacement: 'api_requests_total'
# Drop labels
- regex: 'temporary_.*'
action: labeldrop
# Multiple scrape intervals
scrape_configs:
- job_name: 'fast-metrics'
scrape_interval: 5s
static_configs:
- targets: ['critical-service:9090']
- job_name: 'slow-metrics'
scrape_interval: 60s
static_configs:
- targets: ['batch-service:9090']
Service Discovery
Prometheus can automatically discover scrape targets from various sources, eliminating manual target configuration.
Key Concepts
flowchart TB
subgraph "Service Discovery Sources"
A[Kubernetes API]
B[Consul]
C[EC2/GCE/Azure]
D[DNS SRV Records]
E[File SD]
end
A --> F[Prometheus]
B --> F
C --> F
D --> F
E --> F
F -->|relabel_configs| G[Filtered Targets]
G --> H[Scrape]
Common SD mechanisms:
- kubernetes_sd: Discover pods, services, endpoints, nodes
- consul_sd: Discover services from Consul
- ec2_sd/gce_sd/azure_sd: Cloud provider instances
- dns_sd: DNS-based service discovery
- file_sd: File-based target configuration
Common Patterns
# Kubernetes service discovery
scrape_configs:
# Discover Kubernetes pods
- job_name: 'kubernetes-pods'
kubernetes_sd_configs:
- role: pod
namespaces:
names:
- production
- staging
relabel_configs:
# Only scrape pods with prometheus.io/scrape annotation
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
# Use prometheus.io/path annotation for metrics path
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path]
action: replace
target_label: __metrics_path__
regex: (.+)
# Use prometheus.io/port annotation
- source_labels: [__address__, __meta_kubernetes_pod_annotation_prometheus_io_port]
action: replace
regex: ([^:]+)(?::\d+)?;(\d+)
replacement: $1:$2
target_label: __address__
# Add namespace label
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
# Add pod name label
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
# Discover Kubernetes services
- job_name: 'kubernetes-services'
kubernetes_sd_configs:
- role: service
relabel_configs:
- source_labels: [__meta_kubernetes_service_annotation_prometheus_io_scrape]
action: keep
regex: true
- source_labels: [__meta_kubernetes_service_name]
target_label: service
Examples
# Consul service discovery
scrape_configs:
- job_name: 'consul-services'
consul_sd_configs:
- server: 'consul.example.com:8500'
services: [] # Empty = all services
tags:
- 'prometheus'
relabel_configs:
- source_labels: [__meta_consul_service]
target_label: service
- source_labels: [__meta_consul_dc]
target_label: datacenter
# EC2 service discovery
scrape_configs:
- job_name: 'ec2-instances'
ec2_sd_configs:
- region: eu-west-1
port: 9100
filters:
- name: tag:Environment
values: ['production']
- name: instance-state-name
values: ['running']
relabel_configs:
- source_labels: [__meta_ec2_tag_Name]
target_label: instance_name
- source_labels: [__meta_ec2_availability_zone]
target_label: availability_zone
# File-based service discovery
scrape_configs:
- job_name: 'file-sd'
file_sd_configs:
- files:
- '/etc/prometheus/targets/*.json'
refresh_interval: 5m
# targets/app.json
[
{
"targets": ["app1:9090", "app2:9090"],
"labels": {
"env": "production",
"team": "backend"
}
}
]
# DNS service discovery
scrape_configs:
- job_name: 'dns-sd'
dns_sd_configs:
- names:
- '_prometheus._tcp.example.com'
type: SRV
refresh_interval: 30s
Alerting Rules
Alerting rules define conditions that trigger alerts when met, sending notifications through Alertmanager.
Key Concepts
sequenceDiagram
participant P as Prometheus
participant A as Alertmanager
participant N as Notification Channel
P->>P: Evaluate alert rule
P->>P: Alert fires (condition true)
P->>P: Wait for 'for' duration
P->>A: Send alert
A->>A: Group alerts
A->>A: Apply routes
A->>A: Silence/Inhibit check
A->>N: Send notification
Alert states:
- Inactive: Condition is false
- Pending: Condition is true, waiting for
forduration - Firing: Condition true for
forduration, alert sent
Common Patterns
# alerts.yml
groups:
- name: instance-alerts
rules:
# Instance down alert
- alert: InstanceDown
expr: up == 0
for: 5m
labels:
severity: critical
team: infrastructure
annotations:
summary: "Instance {{ $labels.instance }} is down"
description: "{{ $labels.instance }} of job {{ $labels.job }} has been down for more than 5 minutes."
runbook_url: "https://wiki.example.com/runbooks/instance-down"
# High CPU usage
- alert: HighCPUUsage
expr: |
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 10m
labels:
severity: warning
annotations:
summary: "High CPU usage on {{ $labels.instance }}"
description: "CPU usage is {{ printf \"%.2f\" $value }}%"
# Low disk space
- alert: LowDiskSpace
expr: |
(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
/ node_filesystem_size_bytes * 100) < 10
for: 15m
labels:
severity: warning
annotations:
summary: "Low disk space on {{ $labels.instance }}"
description: "Filesystem {{ $labels.mountpoint }} has {{ printf \"%.2f\" $value }}% free space"
- name: application-alerts
rules:
# High error rate
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/ sum(rate(http_requests_total[5m])) by (service) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "High error rate for {{ $labels.service }}"
description: "Error rate is {{ $value | humanizePercentage }}"
# High latency
- alert: HighLatency
expr: |
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
) > 1
for: 5m
labels:
severity: warning
annotations:
summary: "High latency for {{ $labels.service }}"
description: "95th percentile latency is {{ printf \"%.3f\" $value }}s"
Examples
# Advanced alerting patterns
groups:
- name: slo-alerts
rules:
# Error budget burn rate (Multi-window, multi-burn-rate)
- alert: ErrorBudgetBurn
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/ sum(rate(http_requests_total[1h]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
) > (14.4 * 0.001)
for: 2m
labels:
severity: critical
annotations:
summary: "High error budget burn rate"
description: "Error budget is being consumed at 14.4x the target rate"
# Absent metrics (detect missing scrapes)
- alert: MetricAbsent
expr: absent(up{job="critical-service"})
for: 5m
labels:
severity: critical
annotations:
summary: "Critical service metrics missing"
description: "No metrics received from critical-service for 5 minutes"
# Prediction-based alert
- alert: DiskWillFillIn24Hours
expr: |
predict_linear(node_filesystem_avail_bytes[6h], 24*3600) < 0
for: 1h
labels:
severity: warning
annotations:
summary: "Disk will fill within 24 hours"
description: "Based on current trends, {{ $labels.mountpoint }} will be full in less than 24 hours"
Recording Rules
Recording rules pre-compute frequently used or computationally expensive expressions, storing results as new time series.
Key Concepts
Benefits of recording rules:
- Reduce query latency for dashboards
- Compute expensive aggregations once
- Create derived metrics for simpler queries
- Enable federation of pre-aggregated data
Naming convention: level:metric:operations
- level: aggregation level (e.g., job, instance)
- metric: metric name
- operations: list of functions applied
Common Patterns
# rules/recording-rules.yml
groups:
- name: request-recording-rules
interval: 30s # Override default evaluation interval
rules:
# Pre-compute request rate by job
- record: job:http_requests_total:rate5m
expr: sum(rate(http_requests_total[5m])) by (job)
# Pre-compute error rate by job
- record: job:http_requests_errors:rate5m
expr: sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
# Pre-compute error ratio
- record: job:http_requests_error_ratio:rate5m
expr: |
job:http_requests_errors:rate5m
/ job:http_requests_total:rate5m
# Pre-compute latency percentiles
- record: job:http_request_duration_seconds:p50
expr: |
histogram_quantile(0.50,
sum(rate(http_request_duration_seconds_bucket[5m])) by (job, le)
)
- record: job:http_request_duration_seconds:p95
expr: |
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (job, le)
)
- record: job:http_request_duration_seconds:p99
expr: |
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (job, le)
)
- name: node-recording-rules
rules:
# CPU usage by instance
- record: instance:node_cpu_utilisation:rate5m
expr: |
1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))
# Memory usage by instance
- record: instance:node_memory_utilisation:ratio
expr: |
1 - (
node_memory_MemAvailable_bytes
/ node_memory_MemTotal_bytes
)
# Disk usage by instance and device
- record: instance:node_filesystem_utilisation:ratio
expr: |
1 - (
node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
/ node_filesystem_size_bytes
)
Examples
# Aggregation hierarchy for federation
groups:
- name: aggregation-rules
rules:
# Instance level
- record: instance:http_requests:rate5m
expr: sum(rate(http_requests_total[5m])) by (instance, job)
# Job level (aggregates instance level)
- record: job:http_requests:rate5m
expr: sum(instance:http_requests:rate5m) by (job)
# Cluster level (aggregates job level)
- record: cluster:http_requests:rate5m
expr: sum(job:http_requests:rate5m)
# SLI recording rules for SLO monitoring
- record: sli:http_requests_availability:ratio_rate5m
expr: |
sum(rate(http_requests_total{status!~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
labels:
slo: "availability"
- record: sli:http_requests_latency:ratio_rate5m
expr: |
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/ sum(rate(http_request_duration_seconds_count[5m]))
labels:
slo: "latency"
Common Queries for Monitoring
Essential PromQL queries for monitoring infrastructure and applications.
Infrastructure Monitoring
# CPU
# CPU usage percentage per instance
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# CPU usage by mode
sum by (mode) (rate(node_cpu_seconds_total[5m])) * 100
# Memory
# Memory usage percentage
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100
# Memory breakdown (cache, buffers, used)
node_memory_MemTotal_bytes - node_memory_MemFree_bytes - node_memory_Buffers_bytes - node_memory_Cached_bytes
# Disk
# Disk usage percentage by mount point
(1 - (node_filesystem_avail_bytes / node_filesystem_size_bytes)) * 100
# Disk I/O rate
rate(node_disk_read_bytes_total[5m]) + rate(node_disk_written_bytes_total[5m])
# Network
# Network throughput
rate(node_network_receive_bytes_total[5m]) + rate(node_network_transmit_bytes_total[5m])
# Network errors
rate(node_network_receive_errs_total[5m]) + rate(node_network_transmit_errs_total[5m])
Application Monitoring
# Request rate
# Requests per second by endpoint
sum(rate(http_requests_total[5m])) by (endpoint)
# Request rate trend (compare to 1 hour ago)
sum(rate(http_requests_total[5m])) - sum(rate(http_requests_total[5m] offset 1h))
# Error rate
# Error percentage
sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) * 100
# Errors by status code
sum(rate(http_requests_total{status=~"[45].."}[5m])) by (status)
# Latency
# Average latency
rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m])
# Latency percentiles from histogram
histogram_quantile(0.50, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
# Apdex score (satisfied < 0.3s, tolerating < 1.2s)
(
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m])) +
sum(rate(http_request_duration_seconds_bucket{le="1.2"}[5m]))
) / 2 / sum(rate(http_request_duration_seconds_count[5m]))
Kubernetes Monitoring
# Pod status
# Number of pods by phase
count(kube_pod_status_phase{phase="Running"}) by (namespace)
# Pods not ready
kube_pod_status_ready{condition="false"}
# Container restarts in last hour
increase(kube_pod_container_status_restarts_total[1h]) > 0
# Resource usage
# CPU usage vs requests
sum(rate(container_cpu_usage_seconds_total[5m])) by (pod)
/ sum(kube_pod_container_resource_requests{resource="cpu"}) by (pod)
# Memory usage vs limits
sum(container_memory_working_set_bytes) by (pod)
/ sum(kube_pod_container_resource_limits{resource="memory"}) by (pod)
# Deployment health
# Deployment replicas available vs desired
kube_deployment_status_replicas_available / kube_deployment_spec_replicas
# Failed deployments
kube_deployment_status_replicas_unavailable > 0
Federation Patterns
Federation allows Prometheus servers to scrape metrics from other Prometheus servers, enabling hierarchical or cross-service aggregation.
Key Concepts
graph TB
subgraph "Federated Architecture"
subgraph "Regional Prometheus"
A[Prometheus EU]
B[Prometheus US]
C[Prometheus APAC]
end
D[Global Prometheus] -->|/federate| A
D -->|/federate| B
D -->|/federate| C
D --> E[Global Grafana]
A --> F[EU Grafana]
B --> G[US Grafana]
end
Federation use cases:
- Hierarchical: Aggregate regional data to global view
- Cross-service: Pull specific metrics from another team's Prometheus
- HA Pairs: Synchronise recording rules across replicas
Common Patterns
# Global Prometheus configuration
scrape_configs:
# Federate from regional Prometheus instances
- job_name: 'federate-eu'
honor_labels: true
metrics_path: '/federate'
params:
'match[]':
# Pull pre-aggregated recording rules
- '{__name__=~"job:.*"}'
# Pull specific high-level metrics
- '{__name__=~"cluster:.*"}'
# Pull critical raw metrics
- 'up{job="critical-service"}'
static_configs:
- targets:
- 'prometheus-eu.example.com:9090'
labels:
region: 'eu'
- job_name: 'federate-us'
honor_labels: true
metrics_path: '/federate'
params:
'match[]':
- '{__name__=~"job:.*"}'
- '{__name__=~"cluster:.*"}'
static_configs:
- targets:
- 'prometheus-us.example.com:9090'
labels:
region: 'us'
- job_name: 'federate-apac'
honor_labels: true
metrics_path: '/federate'
params:
'match[]':
- '{__name__=~"job:.*"}'
- '{__name__=~"cluster:.*"}'
static_configs:
- targets:
- 'prometheus-apac.example.com:9090'
labels:
region: 'apac'
Examples
# Cross-service federation
# Pull metrics from another team's Prometheus
scrape_configs:
- job_name: 'federate-payments'
honor_labels: true
metrics_path: '/federate'
params:
'match[]':
# Only pull metrics you need
- 'payment_transactions_total'
- 'payment_latency_seconds_bucket'
static_configs:
- targets:
- 'prometheus-payments.internal:9090'
# Recording rules for federation efficiency
# On regional Prometheus instances, create pre-aggregated metrics
groups:
- name: federation-rules
rules:
- record: region:http_requests:rate5m
expr: sum(rate(http_requests_total[5m]))
labels:
region: "eu"
- record: region:http_errors:rate5m
expr: sum(rate(http_requests_total{status=~"5.."}[5m]))
labels:
region: "eu"
- record: region:http_latency_seconds:p99
expr: |
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
labels:
region: "eu"
# Thanos/Cortex alternative for long-term storage
# Remote write configuration
remote_write:
- url: "http://thanos-receive.monitoring:19291/api/v1/receive"
queue_config:
max_samples_per_send: 10000
batch_send_deadline: 5s
capacity: 50000
Quick Reference
Essential PromQL Functions
| Function | Description | Example |
|---|---|---|
rate() |
Per-second average rate of increase | rate(http_requests_total[5m]) |
irate() |
Instant rate (last two samples) | irate(http_requests_total[5m]) |
increase() |
Total increase over time range | increase(http_requests_total[1h]) |
sum() |
Sum across dimensions | sum(rate(requests[5m])) by (job) |
avg() |
Average across dimensions | avg(temperature) by (location) |
histogram_quantile() |
Calculate percentile from histogram | histogram_quantile(0.95, sum(rate(bucket[5m])) by (le)) |
predict_linear() |
Predict future value | predict_linear(disk_free[1h], 3600*24) |
absent() |
Returns 1 if no series exist | absent(up{job="api"}) |
label_replace() |
Modify labels with regex | label_replace(up, "host", "$1", "instance", "(.*):.*") |
topk() |
Top K elements by value | topk(5, rate(requests[5m])) |
Common Aggregation Patterns
| Pattern | Description |
|---|---|
sum by (label) |
Sum grouped by label |
avg without (label) |
Average excluding label |
count by (label) |
Count series per label value |
max by (label) |
Maximum per group |
min by (label) |
Minimum per group |
stddev by (label) |
Standard deviation per group |
quantile(0.95, metric) |
95th percentile across series |
Prometheus CLI Commands
# Check configuration file
promtool check config /etc/prometheus/prometheus.yml
# Check rules file
promtool check rules /etc/prometheus/rules/*.yml
# Unit test rules
promtool test rules test.yml
# Query Prometheus API
curl -G 'http://localhost:9090/api/v1/query' \
--data-urlencode 'query=up'
# Query range
curl -G 'http://localhost:9090/api/v1/query_range' \
--data-urlencode 'query=rate(http_requests_total[5m])' \
--data-urlencode 'start=2024-01-01T00:00:00Z' \
--data-urlencode 'end=2024-01-01T01:00:00Z' \
--data-urlencode 'step=60s'
# Reload configuration (requires --web.enable-lifecycle)
curl -X POST http://localhost:9090/-/reload
# Check targets
curl http://localhost:9090/api/v1/targets
# Check alerts
curl http://localhost:9090/api/v1/alerts
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Target shows as DOWN | Network/firewall blocking | Check network connectivity and firewall rules. Verify target is exposing metrics on expected port |
| High cardinality warnings | Too many unique label combinations | Remove or reduce high-cardinality labels (user IDs, request IDs). Use recording rules to pre-aggregate |
| Gaps in metrics | Scrape timeouts or target overload | Increase scrape_timeout, check target health, reduce metric count |
| "query processing would load too many samples" | Query touches too much data | Add filters, reduce time range, use recording rules for aggregation |
| Federation pulling too much data | Matching too many series | Use specific match patterns, pre-aggregate with recording rules |
| Alert not firing | for duration not met or expression wrong |
Check alert state in UI, verify expression returns results, check for duration |
| Duplicate alerts | Multiple Prometheus instances | Use Alertmanager deduplication, configure external_labels |
| Counter resets unexpectedly | Application restart or counter bug | Use rate() or increase() which handle resets automatically |
| Histogram quantile returning NaN | No data in buckets or division by zero | Ensure histogram has data, check bucket configuration |
| Slow queries | Large time ranges, missing recording rules | Create recording rules for expensive queries, reduce query scope |
| Out of memory | Too many time series, long retention | Reduce cardinality, adjust retention, add more memory, use remote storage |
| Scrape taking too long | Target returning too many metrics | Filter metrics at source or with metric_relabel_configs, increase timeout |
Debugging Tips
# Check Prometheus logs
journalctl -u prometheus -f
# Verify target is reachable
curl http://target:port/metrics
# Check metric format
curl -s http://target:port/metrics | promtool check metrics
# Debug relabelling
# Add to scrape config temporarily:
# relabel_configs:
# - action: labelmap
# regex: __meta_(.*)
# Check TSDB stats
curl http://localhost:9090/api/v1/status/tsdb
# List all metric names
curl -s http://localhost:9090/api/v1/label/__name__/values | jq
# Check cardinality
curl -s 'http://localhost:9090/api/v1/query?query=count({__name__=~".+"})' | jq
Related Topics
The following topics complement Prometheus and would enhance your monitoring and observability capabilities:
- Grafana - Visualisation platform that pairs with Prometheus for dashboards, providing rich querying and alerting UI
- Alertmanager - Handles alert routing, grouping, silencing, and notification channels for Prometheus alerts
- Loki - Log aggregation system from Grafana Labs that uses PromQL-like syntax (LogQL) for querying logs
- OpenTelemetry - Vendor-neutral observability framework for traces, metrics, and logs that can export to Prometheus
- Kubernetes - Container orchestration platform commonly monitored with Prometheus, with native service discovery support
- Thanos/Cortex - Long-term storage and global query solutions for Prometheus in multi-cluster environments