Grafana
Open-source analytics and interactive visualisation platform for monitoring metrics, logs, and traces from multiple data sources.
Grafana
Open-source analytics and interactive visualisation platform for monitoring metrics, logs, and traces from multiple data sources.
Overview
Grafana is a multi-platform open-source analytics and visualisation application that connects to various data sources (Prometheus, Loki, Elasticsearch, InfluxDB, and more) and presents data through customisable dashboards. It provides a powerful query builder, templating system, alerting capabilities, and role-based access control, making it the standard visualisation layer in modern observability stacks.
graph TB
subgraph "Grafana Architecture"
A[Grafana Server] --> B[(SQLite/MySQL/PostgreSQL)]
A --> C[Dashboard Engine]
A --> D[Alerting Engine]
A --> E[Plugin System]
end
subgraph "Data Sources"
F[Prometheus]
G[Loki]
H[Elasticsearch]
I[InfluxDB]
J[CloudWatch]
end
F --> A
G --> A
H --> A
I --> A
J --> A
C --> K[Dashboards]
D --> L[Notification Channels]
E --> M[Panels/Apps]
L --> N[Email/Slack/PagerDuty]
Data Source Configuration
Data sources connect Grafana to your metrics, logs, and trace backends, enabling queries and visualisation.
Key Concepts
| Concept | Description |
|---|---|
| Data Source | Connection configuration to a backend system |
| Default Data Source | Used when no specific source is selected in panels |
| Provisioning | Automated data source setup via configuration files |
| Correlations | Link related data across different sources |
Built-in data sources include: Prometheus, Loki, Elasticsearch, InfluxDB, MySQL, PostgreSQL, CloudWatch, Azure Monitor, Google Cloud Monitoring, Tempo, Jaeger, and Zipkin.
Common Patterns
# Provisioning data source via YAML
# /etc/grafana/provisioning/datasources/prometheus.yaml
apiVersion: 1
datasources:
# Prometheus data source
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
editable: false
jsonData:
httpMethod: POST
timeInterval: 15s
exemplarTraceIdDestinations:
- name: traceID
datasourceUid: tempo
# Loki data source
- name: Loki
type: loki
access: proxy
url: http://loki:3100
jsonData:
maxLines: 1000
derivedFields:
- name: TraceID
matcherRegex: "traceID=(\\w+)"
url: '$${__value.raw}'
datasourceUid: tempo
# Elasticsearch data source
- name: Elasticsearch
type: elasticsearch
access: proxy
url: http://elasticsearch:9200
database: "[logs-]YYYY.MM.DD"
jsonData:
interval: Daily
timeField: "@timestamp"
esVersion: "8.0.0"
logMessageField: message
logLevelField: level
Examples
# CloudWatch data source with IAM role
- name: CloudWatch
type: cloudwatch
jsonData:
authType: keys
defaultRegion: eu-west-1
secureJsonData:
accessKey: "${AWS_ACCESS_KEY_ID}"
secretKey: "${AWS_SECRET_ACCESS_KEY}"
# InfluxDB 2.x data source
- name: InfluxDB
type: influxdb
access: proxy
url: http://influxdb:8086
jsonData:
version: Flux
organization: myorg
defaultBucket: metrics
tlsSkipVerify: false
secureJsonData:
token: "${INFLUXDB_TOKEN}"
# PostgreSQL data source
- name: PostgreSQL
type: postgres
url: postgres:5432
database: grafana
user: grafana
secureJsonData:
password: "${PG_PASSWORD}"
jsonData:
sslmode: require
maxOpenConns: 100
maxIdleConns: 100
connMaxLifetime: 14400
postgresVersion: 1500
timescaledb: false
Via Grafana UI:
- Navigate to Configuration > Data Sources
- Click "Add data source"
- Select the data source type
- Configure connection settings
- Click "Save & Test"
Dashboard Creation and Panels
Dashboards are collections of panels that visualise data from configured data sources.
Key Concepts
flowchart TB
subgraph "Dashboard Structure"
A[Dashboard] --> B[Row 1]
A --> C[Row 2]
B --> D[Panel 1]
B --> E[Panel 2]
C --> F[Panel 3]
C --> G[Panel 4]
end
subgraph "Panel Components"
H[Query] --> I[Transform]
I --> J[Visualisation]
J --> K[Overrides]
end
D --> H
Dashboard elements:
- Panels: Individual visualisation components
- Rows: Collapsible containers for organising panels
- Variables: Dynamic values for filtering data
- Annotations: Event markers on graphs
- Links: Navigation to other dashboards or external URLs
Common Patterns
{
"dashboard": {
"id": null,
"uid": "abc123",
"title": "Application Metrics",
"tags": ["production", "api"],
"timezone": "browser",
"schemaVersion": 42,
"version": 1,
"refresh": "30s",
"time": {
"from": "now-1h",
"to": "now"
},
"panels": [
{
"id": 1,
"title": "Request Rate",
"type": "timeseries",
"gridPos": {
"x": 0,
"y": 0,
"w": 12,
"h": 8
},
"targets": [
{
"expr": "sum(rate(http_requests_total[5m])) by (method)",
"legendFormat": "{{method}}",
"refId": "A"
}
],
"fieldConfig": {
"defaults": {
"unit": "reqps",
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": null, "color": "green" },
{ "value": 100, "color": "yellow" },
{ "value": 500, "color": "red" }
]
}
}
}
}
]
}
}
Schema version 42 is what current Grafana (v12.2+) writes; Grafana migrates
older schemaVersion values automatically on import, so dashboards exported
from earlier versions still load.
Examples
# Create dashboard via API (use a service account token; API keys were
# removed in Grafana 12 in favour of service accounts)
curl -X POST \
-H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
-H "Content-Type: application/json" \
-d @dashboard.json \
http://localhost:3000/api/dashboards/db
# Export dashboard
curl -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
http://localhost:3000/api/dashboards/uid/abc123 | jq '.dashboard' > exported.json
# Import dashboard (folderUid replaces the deprecated folderId)
curl -X POST \
-H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"dashboard": '"$(cat exported.json)"',
"overwrite": true,
"folderUid": ""
}' \
http://localhost:3000/api/dashboards/db
Panel creation workflow:
- Click "Add panel" or drag from panel library
- Select visualisation type
- Configure data source and query
- Apply transformations if needed
- Configure panel options and field overrides
- Set thresholds and value mappings
- Save dashboard
Query Builders for Different Data Sources
Each data source has a specific query language and builder interface for fetching data.
Key Concepts
| Data Source | Query Language | Key Features |
|---|---|---|
| Prometheus | PromQL | Metrics, instant/range vectors, aggregations |
| Loki | LogQL | Logs, label filtering, line parsing |
| Elasticsearch | Lucene/PPL | Full-text search, aggregations |
| InfluxDB | Flux/InfluxQL | Time series, buckets, transformations |
| MySQL/PostgreSQL | SQL | Tables, joins, aggregations |
Common Patterns
Prometheus queries:
# Request rate by endpoint
sum(rate(http_requests_total{job="api"}[5m])) by (endpoint)
# Error rate percentage
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) * 100
# 95th percentile latency
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# Memory usage percentage
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100
Loki queries:
# Filter logs by label and content
{job="api", namespace="production"} |= "error"
# Parse JSON and filter
{job="api"} | json | level="error" | line_format "{{.message}}"
# Count errors over time
sum(rate({job="api"} |= "error" [5m])) by (pod)
# Extract metrics from logs
sum by (status) (
count_over_time({job="nginx"} | pattern `<_> - - <_> "<_> <_> <_>" <status> <_>` [5m])
)
Elasticsearch queries:
// Lucene query syntax
message: "error" AND level: "error" AND NOT host: "test-*"
// Using query builder
{
"query": "level:error",
"metrics": [
{ "type": "count", "id": "1" },
{ "type": "avg", "field": "response_time", "id": "2" }
],
"bucketAggs": [
{ "type": "date_histogram", "field": "@timestamp", "id": "3" }
]
}
Examples
-- PostgreSQL/MySQL time series query
SELECT
$__timeGroupAlias(created_at, '1m'),
count(*) as "requests",
avg(response_time) as "avg_latency"
FROM requests
WHERE
$__timeFilter(created_at)
AND status_code >= 500
GROUP BY 1
ORDER BY 1
-- Table format query
SELECT
endpoint,
count(*) as total,
avg(response_time) as avg_ms,
max(response_time) as max_ms
FROM requests
WHERE $__timeFilter(created_at)
GROUP BY endpoint
ORDER BY total DESC
LIMIT 10
Query builder features:
- Code mode: Write raw queries
- Builder mode: Visual query construction
- Explain: Show query execution plan
- Inspector: View raw query results and timing
Variables and Templating
Variables enable dynamic dashboards that can be reused across different environments and contexts.
Key Concepts
flowchart LR
A[Variable Definition] --> B{Variable Type}
B --> C[Query]
B --> D[Custom]
B --> E[Constant]
B --> F[Data source]
B --> G[Interval]
B --> H[Text box]
C --> I[Panel Query]
D --> I
E --> I
F --> I
G --> I
H --> I
| Variable Type | Description | Use Case |
|---|---|---|
| Query | Dynamically fetched from data source | List namespaces, instances, jobs |
| Custom | Static comma-separated values | Environment selection |
| Constant | Single hidden value | Dashboard versioning |
| Data source | Select from available sources | Multi-cluster dashboards |
| Interval | Time interval selection | Query resolution control |
| Text box | Free-form text input | Search filters |
| Ad hoc filters | Dynamic key-value filters | Exploratory filtering |
Common Patterns
# Provisioning variables (within dashboard JSON)
"templating": {
"list": [
{
"name": "datasource",
"type": "datasource",
"query": "prometheus",
"current": {
"text": "Prometheus",
"value": "Prometheus"
},
"hide": 0
},
{
"name": "namespace",
"type": "query",
"datasource": "${datasource}",
"query": "label_values(kube_pod_info, namespace)",
"refresh": 2,
"sort": 1,
"multi": true,
"includeAll": true,
"allValue": ".*"
},
{
"name": "pod",
"type": "query",
"datasource": "${datasource}",
"query": "label_values(kube_pod_info{namespace=~\"$namespace\"}, pod)",
"refresh": 2,
"sort": 1,
"multi": true,
"includeAll": true
},
{
"name": "interval",
"type": "interval",
"query": "1m,5m,10m,30m,1h",
"current": {
"text": "5m",
"value": "5m"
},
"auto": true,
"auto_min": "10s",
"auto_count": 100
}
]
}
Examples
Using variables in queries:
# Prometheus query with variables
sum(rate(http_requests_total{namespace=~"$namespace", pod=~"$pod"}[$interval])) by (method)
# With All option using regex
sum(rate(container_cpu_usage_seconds_total{namespace=~"$namespace"}[5m])) by (pod)
# Loki query with variables
{namespace="$namespace", pod=~"$pod"} |= "$search"
-- SQL query with variables
SELECT * FROM metrics
WHERE namespace IN ($namespace)
AND $__timeFilter(timestamp)
Variable syntax:
$variable- Simple substitution${variable}- Explicit boundaries${variable:csv}- Format as comma-separated values${variable:pipe}- Format as pipe-separated values${variable:regex}- Format as regex (escaped)${variable:raw}- Raw value without escaping
Chained variables example:
$cluster- Query:label_values(up, cluster)$namespace- Query:label_values(kube_pod_info{cluster="$cluster"}, namespace)$deployment- Query:label_values(kube_deployment_labels{cluster="$cluster", namespace="$namespace"}, deployment)
Alerting Setup
Grafana Alerting evaluates alert rules and sends notifications when conditions are met.
Key Concepts
sequenceDiagram
participant AR as Alert Rule
participant AE as Alert Engine
participant AM as Alertmanager
participant CP as Contact Point
participant NP as Notification Policy
participant RC as Receiver
AR->>AE: Evaluate query
AE->>AE: Check condition
AE->>AM: Send alert
AM->>NP: Route alert
NP->>CP: Match labels
CP->>RC: Send notification
| Component | Description |
|---|---|
| Alert Rule | Query and condition that triggers alert |
| Contact Point | Notification destination (email, Slack, PagerDuty) |
| Notification Policy | Routing rules based on labels |
| Silence | Temporarily suppress notifications |
| Mute Timing | Recurring periods to suppress notifications |
Common Patterns
# Provisioning alert rules
# /etc/grafana/provisioning/alerting/alerts.yaml
apiVersion: 1
groups:
- orgId: 1
name: infrastructure
folder: alerts
interval: 1m
rules:
- uid: high-cpu-usage
title: High CPU Usage
condition: C
data:
- refId: A
relativeTimeRange:
from: 300
to: 0
datasourceUid: prometheus
model:
expr: 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
instant: false
intervalMs: 1000
maxDataPoints: 43200
refId: A
- refId: B
relativeTimeRange:
from: 300
to: 0
datasourceUid: __expr__
model:
conditions:
- evaluator:
params: [80]
type: gt
operator:
type: and
query:
params: [A]
reducer:
type: last
refId: B
type: classic_conditions
- refId: C
datasourceUid: __expr__
model:
expression: B
refId: C
type: threshold
for: 5m
labels:
severity: warning
team: infrastructure
annotations:
summary: "High CPU usage on {{ $labels.instance }}"
description: "CPU usage is {{ $values.A }}%"
runbook_url: https://wiki.example.com/runbooks/high-cpu
# Provisioning contact points
contactPoints:
- orgId: 1
name: slack-infrastructure
receivers:
- uid: slack-infra
type: slack
settings:
url: "${SLACK_WEBHOOK_URL}"
recipient: "#infrastructure-alerts"
title: |
{{ template "slack.default.title" . }}
text: |
{{ template "slack.default.text" . }}
# Provisioning notification policies
policies:
- orgId: 1
receiver: slack-infrastructure
group_by:
- alertname
- severity
routes:
- receiver: pagerduty-critical
matchers:
- severity = critical
continue: true
- receiver: slack-infrastructure
matchers:
- team = infrastructure
Examples
Alert rule expressions:
# Multi-condition alert
data:
- refId: A
expr: sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
- refId: B
expr: sum(rate(http_requests_total[5m])) by (service)
- refId: C
datasourceUid: __expr__
model:
expression: A / B * 100
type: math
- refId: D
datasourceUid: __expr__
model:
expression: C
conditions:
- evaluator:
type: gt
params: [5]
type: threshold
condition: D
Contact point configurations:
# Email contact point
- type: email
settings:
addresses: oncall@example.com
singleEmail: true
# PagerDuty contact point
- type: pagerduty
settings:
integrationKey: "${PAGERDUTY_KEY}"
severity: critical
class: infrastructure
component: grafana
# Webhook contact point
- type: webhook
settings:
url: https://api.example.com/alerts
httpMethod: POST
username: grafana
password: "${WEBHOOK_PASSWORD}"
Creating alerts via UI:
- Navigate to Alerting > Alert rules
- Click "New alert rule"
- Define query and conditions
- Set evaluation behaviour (folder, group, interval)
- Add labels and annotations
- Configure notification policy or contact point
Common Visualisation Types
Grafana provides numerous panel types for different data visualisation needs.
Key Concepts
| Visualisation | Best For | Data Type |
|---|---|---|
| Time series | Metrics over time | Time series |
| Stat | Single values, KPIs | Scalar |
| Gauge | Progress/percentage | Scalar |
| Bar chart | Comparisons | Categorical |
| Table | Detailed data | Tabular |
| Heatmap | Distribution over time | Histogram |
| Logs | Log entries | Log streams |
| Node graph | Relationships | Graph data |
| Geomap | Geographical data | Coordinates |
Common Patterns
Time series panel:
{
"type": "timeseries",
"fieldConfig": {
"defaults": {
"custom": {
"drawStyle": "line",
"lineInterpolation": "smooth",
"fillOpacity": 10,
"gradientMode": "scheme",
"showPoints": "auto",
"pointSize": 5,
"stacking": { "mode": "none" },
"axisPlacement": "auto",
"spanNulls": true
},
"color": { "mode": "palette-classic" },
"unit": "bytes",
"decimals": 2,
"min": 0
}
},
"options": {
"legend": {
"displayMode": "table",
"placement": "bottom",
"calcs": ["mean", "max", "last"]
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
}
}
Stat panel:
{
"type": "stat",
"fieldConfig": {
"defaults": {
"unit": "percent",
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": null, "color": "green" },
{ "value": 80, "color": "yellow" },
{ "value": 90, "color": "red" }
]
},
"mappings": [
{
"type": "range",
"options": {
"from": 0,
"to": 50,
"result": { "text": "Low", "color": "green" }
}
}
]
}
},
"options": {
"reduceOptions": {
"calcs": ["lastNotNull"],
"fields": "",
"values": false
},
"orientation": "horizontal",
"textMode": "auto",
"colorMode": "background",
"graphMode": "area",
"justifyMode": "center"
}
}
Examples
Table panel with transformations:
{
"type": "table",
"transformations": [
{
"id": "organize",
"options": {
"excludeByName": { "Time": true },
"renameByName": {
"instance": "Host",
"Value": "CPU %"
}
}
},
{
"id": "sortBy",
"options": {
"fields": {},
"sort": [{ "field": "CPU %", "desc": true }]
}
}
],
"fieldConfig": {
"overrides": [
{
"matcher": { "id": "byName", "options": "CPU %" },
"properties": [
{ "id": "custom.cellOptions", "value": { "type": "color-background" } },
{ "id": "thresholds", "value": {
"mode": "absolute",
"steps": [
{ "value": null, "color": "green" },
{ "value": 70, "color": "yellow" },
{ "value": 90, "color": "red" }
]
}}
]
}
]
}
}
Heatmap panel:
{
"type": "heatmap",
"options": {
"calculate": false,
"yAxis": {
"axisPlacement": "left",
"unit": "s"
},
"cellGap": 1,
"color": {
"scheme": "Spectral",
"mode": "scheme",
"fill": "dark-orange",
"scale": "exponential",
"exponent": 0.5
},
"tooltip": {
"show": true,
"yHistogram": true
},
"legend": { "show": true }
}
}
Bar gauge panel:
{
"type": "bargauge",
"options": {
"reduceOptions": {
"calcs": ["lastNotNull"]
},
"orientation": "horizontal",
"displayMode": "gradient",
"showUnfilled": true,
"minVizWidth": 0,
"minVizHeight": 10
}
}
Dashboard Organisation and Folders
Organise dashboards using folders, tags, and proper naming conventions for maintainability.
Key Concepts
| Concept | Description |
|---|---|
| Folders | Container for grouping related dashboards |
| Tags | Searchable labels applied to dashboards |
| Permissions | Role-based access at folder and dashboard level |
| Starring | Bookmark frequently used dashboards |
| Playlists | Rotating display of dashboards |
Common Patterns
# Folder provisioning
# /etc/grafana/provisioning/dashboards/default.yaml
apiVersion: 1
providers:
- name: 'infrastructure'
orgId: 1
folder: 'Infrastructure'
folderUid: 'infrastructure'
type: file
disableDeletion: false
updateIntervalSeconds: 30
allowUiUpdates: true
options:
path: /var/lib/grafana/dashboards/infrastructure
- name: 'applications'
orgId: 1
folder: 'Applications'
folderUid: 'applications'
type: file
options:
path: /var/lib/grafana/dashboards/applications
- name: 'business'
orgId: 1
folder: 'Business Metrics'
folderUid: 'business'
type: file
options:
path: /var/lib/grafana/dashboards/business
Recommended folder structure:
dashboards/
├── infrastructure/
│ ├── node-overview.json
│ ├── kubernetes-cluster.json
│ └── network-monitoring.json
├── applications/
│ ├── api-service.json
│ ├── web-frontend.json
│ └── worker-service.json
├── databases/
│ ├── postgresql.json
│ ├── redis.json
│ └── elasticsearch.json
├── business/
│ ├── revenue-metrics.json
│ └── user-engagement.json
└── sre/
├── slo-dashboard.json
└── incident-response.json
Examples
# Create folder via API
curl -X POST \
-H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
-H "Content-Type: application/json" \
-d '{"title": "Production"}' \
http://localhost:3000/api/folders
# List folders
curl -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
http://localhost:3000/api/folders
# Set folder permissions
curl -X POST \
-H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"items": [
{"role": "Viewer", "permission": 1},
{"role": "Editor", "permission": 2},
{"teamId": 1, "permission": 4}
]
}' \
http://localhost:3000/api/folders/abc123/permissions
# Search dashboards by tag
curl -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
"http://localhost:3000/api/search?tag=production&type=dash-db"
# Create playlist
curl -X POST \
-H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"name": "NOC Display",
"interval": "5m",
"items": [
{"type": "dashboard_by_uid", "value": "abc123"},
{"type": "dashboard_by_uid", "value": "def456"}
]
}' \
http://localhost:3000/api/playlists
Tagging conventions:
- Environment:
production,staging,development - Team:
platform,backend,frontend,sre - Service type:
api,database,cache,queue - Alert level:
critical,warning,info
Provisioning Dashboards
Automate dashboard deployment using Grafana's provisioning system for version control and consistency.
Key Concepts
flowchart TB
subgraph "Provisioning Flow"
A[Git Repository] --> B[CI/CD Pipeline]
B --> C[Deploy to Grafana]
C --> D["/etc/grafana/provisioning"]
D --> E{Dashboard Provider}
E --> F[(Grafana Database)]
end
subgraph "Provisioning Types"
G[Dashboards]
H[Data Sources]
I[Alert Rules]
J[Contact Points]
K[Notification Policies]
end
D --> G
D --> H
D --> I
D --> J
D --> K
Provisioning directory structure:
/etc/grafana/provisioning/
├── dashboards/
│ ├── default.yaml # Provider configuration
│ └── dashboards/ # Dashboard JSON files
├── datasources/
│ └── datasources.yaml # Data source configurations
├── alerting/
│ ├── alerts.yaml # Alert rules
│ └── contactpoints.yaml # Contact points
├── notifiers/ # Legacy alerting (deprecated)
└── plugins/
└── plugins.yaml # Plugin installations
Common Patterns
# Dashboard provider configuration
# /etc/grafana/provisioning/dashboards/default.yaml
apiVersion: 1
providers:
- name: 'default'
orgId: 1
folder: ''
folderUid: ''
type: file
disableDeletion: false
updateIntervalSeconds: 10
allowUiUpdates: false
options:
path: /var/lib/grafana/dashboards
foldersFromFilesStructure: true
# Data source provisioning
# /etc/grafana/provisioning/datasources/datasources.yaml
apiVersion: 1
deleteDatasources:
- name: OldPrometheus
orgId: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
orgId: 1
url: http://prometheus:9090
isDefault: true
version: 1
editable: false
jsonData:
httpMethod: POST
manageAlerts: true
prometheusType: Prometheus
prometheusVersion: 2.40.0
Examples
Dockerfile with provisioning:
FROM grafana/grafana:12.0.0
# Copy provisioning configurations
COPY provisioning/ /etc/grafana/provisioning/
# Copy dashboards
COPY dashboards/ /var/lib/grafana/dashboards/
# Copy plugins
COPY plugins/ /var/lib/grafana/plugins/
# Set environment variables
ENV GF_SECURITY_ADMIN_PASSWORD=admin \
GF_USERS_ALLOW_SIGN_UP=false \
GF_AUTH_ANONYMOUS_ENABLED=true \
GF_AUTH_ANONYMOUS_ORG_ROLE=Viewer
Kubernetes ConfigMap for provisioning:
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-dashboards-config
namespace: monitoring
data:
dashboards.yaml: |
apiVersion: 1
providers:
- name: 'default'
orgId: 1
folder: ''
type: file
disableDeletion: false
options:
path: /var/lib/grafana/dashboards
---
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-dashboards
namespace: monitoring
labels:
grafana_dashboard: "1"
data:
api-metrics.json: |
{
"dashboard": {
"title": "API Metrics",
...
}
}
Helm values for Grafana provisioning:
# values.yaml for grafana helm chart
grafana:
dashboardProviders:
dashboardproviders.yaml:
apiVersion: 1
providers:
- name: 'default'
orgId: 1
folder: ''
type: file
disableDeletion: false
editable: true
options:
path: /var/lib/grafana/dashboards/default
dashboards:
default:
node-exporter:
gnetId: 1860
revision: 27
datasource: Prometheus
kubernetes-cluster:
gnetId: 6417
revision: 1
datasource: Prometheus
custom-dashboard:
file: dashboards/custom.json
datasources:
datasources.yaml:
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
url: http://prometheus-server
isDefault: true
alerting:
rules.yaml:
apiVersion: 1
groups:
- orgId: 1
name: my-alerts
folder: alerts
interval: 1m
rules: []
Grafana Terraform provider:
resource "grafana_folder" "production" {
title = "Production"
}
resource "grafana_dashboard" "api_metrics" {
folder = grafana_folder.production.id
config_json = file("dashboards/api-metrics.json")
depends_on = [grafana_data_source.prometheus]
}
resource "grafana_data_source" "prometheus" {
type = "prometheus"
name = "Prometheus"
url = "http://prometheus:9090"
json_data_encoded = jsonencode({
httpMethod = "POST"
})
}
Quick Reference
Grafana CLI Commands
| Command | Description |
|---|---|
grafana-cli plugins list-remote |
List available plugins |
grafana-cli plugins install <plugin> |
Install a plugin |
grafana-cli plugins update-all |
Update all plugins |
grafana-cli admin reset-admin-password <password> |
Reset admin password |
grafana cli admin secrets-migration re-encrypt |
Re-encrypt secrets with the configured encryption (idempotent) |
The old grafana-cli admin data-migration encrypt-datasource-passwords is a
legacy pre-v7 migration (moves plaintext data source passwords into
secure_json_data) and is not needed on modern installs.
API Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/dashboards/db |
POST | Create/update dashboard |
/api/dashboards/uid/:uid |
GET | Get dashboard by UID |
/api/datasources |
GET/POST | List/create data sources |
/api/folders |
GET/POST | List/create folders |
/api/search |
GET | Search dashboards |
/api/v1/provisioning/alert-rules |
GET | List alert rules |
/api/annotations |
GET/POST | Manage annotations |
/api/org/users |
GET | List organisation users |
Common Variable Queries (Prometheus)
| Purpose | Query |
|---|---|
| All namespaces | label_values(kube_pod_info, namespace) |
| Pods in namespace | label_values(kube_pod_info{namespace="$namespace"}, pod) |
| All instances | label_values(up, instance) |
| Jobs | label_values(up, job) |
| Metric names | label_names() |
| Label values | label_values(metric_name, label_name) |
Panel Configuration Options
| Option | Description |
|---|---|
unit |
Display unit (bytes, percent, reqps, etc.) |
decimals |
Number of decimal places |
min/max |
Axis boundaries |
thresholds |
Colour-coded value ranges |
mappings |
Value to text/colour mappings |
overrides |
Field-specific configurations |
transformations |
Data manipulation steps |
Environment Variables
| Variable | Description | Default |
|---|---|---|
GF_SECURITY_ADMIN_USER |
Admin username | admin |
GF_SECURITY_ADMIN_PASSWORD |
Admin password | admin |
GF_SERVER_ROOT_URL |
Full public URL | - |
GF_DATABASE_TYPE |
Database type | sqlite3 |
GF_AUTH_ANONYMOUS_ENABLED |
Allow anonymous access | false |
GF_INSTALL_PLUGINS |
Comma-separated plugin list | - |
GF_PATHS_PROVISIONING |
Provisioning path | /etc/grafana/provisioning |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Data source test fails | Network/auth issues | Verify URL, credentials, and network connectivity. Check Grafana server logs for detailed errors |
| No data in panel | Query returns empty | Use Query Inspector to debug. Check time range matches data. Verify data source is correct |
| Variables not loading | Query syntax error | Check variable query syntax. Ensure data source is selected. Look for errors in variable settings |
| Dashboard not updating | Browser cache | Hard refresh (Ctrl+Shift+R). Check auto-refresh interval. Verify data source is receiving data |
| Alerts not firing | Condition never met | Check alert rule evaluation in Alert Rules page. Verify expression in explore view. Check "for" duration |
| Provisioned dashboard shows "cannot save" | Read-only provisioning | Set allowUiUpdates: true in provider config, or edit source files directly |
| High memory usage | Too many panels/queries | Reduce panel count. Use recording rules for expensive queries. Implement pagination |
| Slow dashboard loading | Complex queries | Pre-aggregate with recording rules. Reduce time range. Use query caching |
| Permission denied | RBAC restrictions | Check user role and team permissions. Verify folder permissions |
| Plugin not found | Not installed | Install via CLI: grafana-cli plugins install <plugin-id> |
| Dashboard UID conflict | Duplicate UIDs | Generate new UID or remove conflicting dashboard before import |
| Template variable recursion | Circular dependency | Review variable dependencies. Ensure no variable references itself indirectly |
Debugging Tips
# Check Grafana logs
journalctl -u grafana-server -f
# Or Docker logs
docker logs grafana -f
# Test data source connectivity
curl -v http://prometheus:9090/api/v1/query?query=up
# Check provisioning status
curl -H "Authorization: Bearer ${SA_TOKEN}" \
http://localhost:3000/api/admin/provisioning/dashboards/reload
# Export dashboard for debugging
curl -H "Authorization: Bearer ${SA_TOKEN}" \
http://localhost:3000/api/dashboards/uid/abc123 | jq
# Check plugin status
curl -H "Authorization: Bearer ${SA_TOKEN}" \
http://localhost:3000/api/plugins
# Verify alert rule status
curl -H "Authorization: Bearer ${SA_TOKEN}" \
http://localhost:3000/api/v1/provisioning/alert-rules
# Database backup (SQLite)
sqlite3 /var/lib/grafana/grafana.db ".backup '/backup/grafana.db'"
Query Inspector usage:
- Open panel in edit mode
- Click "Query Inspector" tab
- View "Query" for sent request
- View "Data" for response
- Check "Stats" for timing information
- Export results as CSV for analysis
Related Topics
The following topics complement Grafana and would enhance your monitoring and observability capabilities:
- Prometheus - Time-series database that pairs with Grafana as the primary metrics data source, providing PromQL for powerful queries
- Loki - Log aggregation system from Grafana Labs that integrates seamlessly with Grafana for log exploration and correlation with metrics
- Alertmanager - Handles alert routing, grouping, and notification dispatch for Prometheus alerts, complementing Grafana's alerting
- OpenTelemetry - Vendor-neutral observability framework for collecting traces, metrics, and logs that can be visualised in Grafana
- Tempo - Grafana's distributed tracing backend that integrates with Grafana for trace visualisation and correlation
- Kubernetes - Container orchestration platform commonly monitored with Grafana dashboards using Prometheus metrics and Loki logs