Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Grafana

Open-source analytics and interactive visualisation platform for monitoring metrics, logs, and traces from multiple data sources.

Grafana

Open-source analytics and interactive visualisation platform for monitoring metrics, logs, and traces from multiple data sources.

Overview

Grafana is a multi-platform open-source analytics and visualisation application that connects to various data sources (Prometheus, Loki, Elasticsearch, InfluxDB, and more) and presents data through customisable dashboards. It provides a powerful query builder, templating system, alerting capabilities, and role-based access control, making it the standard visualisation layer in modern observability stacks.

Data SourcesGrafana ArchitectureGrafana ServerSQLite/MySQL/PostgreSQLDashboard EngineAlerting EnginePlugin SystemPrometheusLokiElasticsearchInfluxDBCloudWatchDashboardsNotificationChannelsPanels/AppsEmail/Slack/PagerDutyData SourcesGrafana ArchitectureGrafana ServerSQLite/MySQL/PostgreSQLDashboard EngineAlerting EnginePlugin SystemPrometheusLokiElasticsearchInfluxDBCloudWatchDashboardsNotificationChannelsPanels/AppsEmail/Slack/PagerDuty

Data Source Configuration

Data sources connect Grafana to your metrics, logs, and trace backends, enabling queries and visualisation.

Key Concepts

Concept Description
Data Source Connection configuration to a backend system
Default Data Source Used when no specific source is selected in panels
Provisioning Automated data source setup via configuration files
Correlations Link related data across different sources

Built-in data sources include: Prometheus, Loki, Elasticsearch, InfluxDB, MySQL, PostgreSQL, CloudWatch, Azure Monitor, Google Cloud Monitoring, Tempo, Jaeger, and Zipkin.

Common Patterns

# Provisioning data source via YAML
# /etc/grafana/provisioning/datasources/prometheus.yaml
apiVersion: 1

datasources:
  # Prometheus data source
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true
    editable: false
    jsonData:
      httpMethod: POST
      timeInterval: 15s
      exemplarTraceIdDestinations:
        - name: traceID
          datasourceUid: tempo

  # Loki data source
  - name: Loki
    type: loki
    access: proxy
    url: http://loki:3100
    jsonData:
      maxLines: 1000
      derivedFields:
        - name: TraceID
          matcherRegex: "traceID=(\\w+)"
          url: '$${__value.raw}'
          datasourceUid: tempo

  # Elasticsearch data source
  - name: Elasticsearch
    type: elasticsearch
    access: proxy
    url: http://elasticsearch:9200
    database: "[logs-]YYYY.MM.DD"
    jsonData:
      interval: Daily
      timeField: "@timestamp"
      esVersion: "8.0.0"
      logMessageField: message
      logLevelField: level

Examples

# CloudWatch data source with IAM role
- name: CloudWatch
  type: cloudwatch
  jsonData:
    authType: keys
    defaultRegion: eu-west-1
  secureJsonData:
    accessKey: "${AWS_ACCESS_KEY_ID}"
    secretKey: "${AWS_SECRET_ACCESS_KEY}"

# InfluxDB 2.x data source
- name: InfluxDB
  type: influxdb
  access: proxy
  url: http://influxdb:8086
  jsonData:
    version: Flux
    organization: myorg
    defaultBucket: metrics
    tlsSkipVerify: false
  secureJsonData:
    token: "${INFLUXDB_TOKEN}"

# PostgreSQL data source
- name: PostgreSQL
  type: postgres
  url: postgres:5432
  database: grafana
  user: grafana
  secureJsonData:
    password: "${PG_PASSWORD}"
  jsonData:
    sslmode: require
    maxOpenConns: 100
    maxIdleConns: 100
    connMaxLifetime: 14400
    postgresVersion: 1500
    timescaledb: false

Via Grafana UI:

  1. Navigate to Configuration > Data Sources
  2. Click "Add data source"
  3. Select the data source type
  4. Configure connection settings
  5. Click "Save & Test"

Dashboard Creation and Panels

Dashboards are collections of panels that visualise data from configured data sources.

Key Concepts

Panel ComponentsDashboard StructureDashboardRow 1Row 2Panel 1Panel 2Panel 3Panel 4QueryTransformVisualisationOverridesPanel ComponentsDashboard StructureDashboardRow 1Row 2Panel 1Panel 2Panel 3Panel 4QueryTransformVisualisationOverrides

Dashboard elements:

  • Panels: Individual visualisation components
  • Rows: Collapsible containers for organising panels
  • Variables: Dynamic values for filtering data
  • Annotations: Event markers on graphs
  • Links: Navigation to other dashboards or external URLs

Common Patterns

{
  "dashboard": {
    "id": null,
    "uid": "abc123",
    "title": "Application Metrics",
    "tags": ["production", "api"],
    "timezone": "browser",
    "schemaVersion": 42,
    "version": 1,
    "refresh": "30s",
    "time": {
      "from": "now-1h",
      "to": "now"
    },
    "panels": [
      {
        "id": 1,
        "title": "Request Rate",
        "type": "timeseries",
        "gridPos": {
          "x": 0,
          "y": 0,
          "w": 12,
          "h": 8
        },
        "targets": [
          {
            "expr": "sum(rate(http_requests_total[5m])) by (method)",
            "legendFormat": "{{method}}",
            "refId": "A"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "reqps",
            "thresholds": {
              "mode": "absolute",
              "steps": [
                { "value": null, "color": "green" },
                { "value": 100, "color": "yellow" },
                { "value": 500, "color": "red" }
              ]
            }
          }
        }
      }
    ]
  }
}

Schema version 42 is what current Grafana (v12.2+) writes; Grafana migrates older schemaVersion values automatically on import, so dashboards exported from earlier versions still load.

Examples

# Create dashboard via API (use a service account token; API keys were
# removed in Grafana 12 in favour of service accounts)
curl -X POST \
  -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  -H "Content-Type: application/json" \
  -d @dashboard.json \
  http://localhost:3000/api/dashboards/db

# Export dashboard
curl -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  http://localhost:3000/api/dashboards/uid/abc123 | jq '.dashboard' > exported.json

# Import dashboard (folderUid replaces the deprecated folderId)
curl -X POST \
  -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "dashboard": '"$(cat exported.json)"',
    "overwrite": true,
    "folderUid": ""
  }' \
  http://localhost:3000/api/dashboards/db

Panel creation workflow:

  1. Click "Add panel" or drag from panel library
  2. Select visualisation type
  3. Configure data source and query
  4. Apply transformations if needed
  5. Configure panel options and field overrides
  6. Set thresholds and value mappings
  7. Save dashboard

Query Builders for Different Data Sources

Each data source has a specific query language and builder interface for fetching data.

Key Concepts

Data Source Query Language Key Features
Prometheus PromQL Metrics, instant/range vectors, aggregations
Loki LogQL Logs, label filtering, line parsing
Elasticsearch Lucene/PPL Full-text search, aggregations
InfluxDB Flux/InfluxQL Time series, buckets, transformations
MySQL/PostgreSQL SQL Tables, joins, aggregations

Common Patterns

Prometheus queries:

# Request rate by endpoint
sum(rate(http_requests_total{job="api"}[5m])) by (endpoint)

# Error rate percentage
sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m])) * 100

# 95th percentile latency
histogram_quantile(0.95,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)

# Memory usage percentage
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100

Loki queries:

# Filter logs by label and content
{job="api", namespace="production"} |= "error"

# Parse JSON and filter
{job="api"} | json | level="error" | line_format "{{.message}}"

# Count errors over time
sum(rate({job="api"} |= "error" [5m])) by (pod)

# Extract metrics from logs
sum by (status) (
  count_over_time({job="nginx"} | pattern `<_> - - <_> "<_> <_> <_>" <status> <_>` [5m])
)

Elasticsearch queries:

// Lucene query syntax
message: "error" AND level: "error" AND NOT host: "test-*"

// Using query builder
{
  "query": "level:error",
  "metrics": [
    { "type": "count", "id": "1" },
    { "type": "avg", "field": "response_time", "id": "2" }
  ],
  "bucketAggs": [
    { "type": "date_histogram", "field": "@timestamp", "id": "3" }
  ]
}

Examples

-- PostgreSQL/MySQL time series query
SELECT
  $__timeGroupAlias(created_at, '1m'),
  count(*) as "requests",
  avg(response_time) as "avg_latency"
FROM requests
WHERE
  $__timeFilter(created_at)
  AND status_code >= 500
GROUP BY 1
ORDER BY 1

-- Table format query
SELECT
  endpoint,
  count(*) as total,
  avg(response_time) as avg_ms,
  max(response_time) as max_ms
FROM requests
WHERE $__timeFilter(created_at)
GROUP BY endpoint
ORDER BY total DESC
LIMIT 10

Query builder features:

  • Code mode: Write raw queries
  • Builder mode: Visual query construction
  • Explain: Show query execution plan
  • Inspector: View raw query results and timing

Variables and Templating

Variables enable dynamic dashboards that can be reused across different environments and contexts.

Key Concepts

Variable DefinitionVariable TypeQueryCustomConstantData sourceIntervalText boxPanel QueryVariable DefinitionVariable TypeQueryCustomConstantData sourceIntervalText boxPanel Query
Variable Type Description Use Case
Query Dynamically fetched from data source List namespaces, instances, jobs
Custom Static comma-separated values Environment selection
Constant Single hidden value Dashboard versioning
Data source Select from available sources Multi-cluster dashboards
Interval Time interval selection Query resolution control
Text box Free-form text input Search filters
Ad hoc filters Dynamic key-value filters Exploratory filtering

Common Patterns

# Provisioning variables (within dashboard JSON)
"templating": {
  "list": [
    {
      "name": "datasource",
      "type": "datasource",
      "query": "prometheus",
      "current": {
        "text": "Prometheus",
        "value": "Prometheus"
      },
      "hide": 0
    },
    {
      "name": "namespace",
      "type": "query",
      "datasource": "${datasource}",
      "query": "label_values(kube_pod_info, namespace)",
      "refresh": 2,
      "sort": 1,
      "multi": true,
      "includeAll": true,
      "allValue": ".*"
    },
    {
      "name": "pod",
      "type": "query",
      "datasource": "${datasource}",
      "query": "label_values(kube_pod_info{namespace=~\"$namespace\"}, pod)",
      "refresh": 2,
      "sort": 1,
      "multi": true,
      "includeAll": true
    },
    {
      "name": "interval",
      "type": "interval",
      "query": "1m,5m,10m,30m,1h",
      "current": {
        "text": "5m",
        "value": "5m"
      },
      "auto": true,
      "auto_min": "10s",
      "auto_count": 100
    }
  ]
}

Examples

Using variables in queries:

# Prometheus query with variables
sum(rate(http_requests_total{namespace=~"$namespace", pod=~"$pod"}[$interval])) by (method)

# With All option using regex
sum(rate(container_cpu_usage_seconds_total{namespace=~"$namespace"}[5m])) by (pod)
# Loki query with variables
{namespace="$namespace", pod=~"$pod"} |= "$search"
-- SQL query with variables
SELECT * FROM metrics
WHERE namespace IN ($namespace)
  AND $__timeFilter(timestamp)

Variable syntax:

  • $variable - Simple substitution
  • ${variable} - Explicit boundaries
  • ${variable:csv} - Format as comma-separated values
  • ${variable:pipe} - Format as pipe-separated values
  • ${variable:regex} - Format as regex (escaped)
  • ${variable:raw} - Raw value without escaping

Chained variables example:

  1. $cluster - Query: label_values(up, cluster)
  2. $namespace - Query: label_values(kube_pod_info{cluster="$cluster"}, namespace)
  3. $deployment - Query: label_values(kube_deployment_labels{cluster="$cluster", namespace="$namespace"}, deployment)

Alerting Setup

Grafana Alerting evaluates alert rules and sends notifications when conditions are met.

Key Concepts

ReceiverNotification PolicyContact PointAlertmanagerAlert EngineAlert RuleReceiverNotification PolicyContact PointAlertmanagerAlert EngineAlert RuleEvaluate queryCheck conditionSend alertRoute alertMatch labelsSend notificationReceiverNotification PolicyContact PointAlertmanagerAlert EngineAlert RuleReceiverNotification PolicyContact PointAlertmanagerAlert EngineAlert RuleEvaluate queryCheck conditionSend alertRoute alertMatch labelsSend notification
Component Description
Alert Rule Query and condition that triggers alert
Contact Point Notification destination (email, Slack, PagerDuty)
Notification Policy Routing rules based on labels
Silence Temporarily suppress notifications
Mute Timing Recurring periods to suppress notifications

Common Patterns

# Provisioning alert rules
# /etc/grafana/provisioning/alerting/alerts.yaml
apiVersion: 1

groups:
  - orgId: 1
    name: infrastructure
    folder: alerts
    interval: 1m
    rules:
      - uid: high-cpu-usage
        title: High CPU Usage
        condition: C
        data:
          - refId: A
            relativeTimeRange:
              from: 300
              to: 0
            datasourceUid: prometheus
            model:
              expr: 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
              instant: false
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: B
            relativeTimeRange:
              from: 300
              to: 0
            datasourceUid: __expr__
            model:
              conditions:
                - evaluator:
                    params: [80]
                    type: gt
                  operator:
                    type: and
                  query:
                    params: [A]
                  reducer:
                    type: last
              refId: B
              type: classic_conditions
          - refId: C
            datasourceUid: __expr__
            model:
              expression: B
              refId: C
              type: threshold
        for: 5m
        labels:
          severity: warning
          team: infrastructure
        annotations:
          summary: "High CPU usage on {{ $labels.instance }}"
          description: "CPU usage is {{ $values.A }}%"
          runbook_url: https://wiki.example.com/runbooks/high-cpu

# Provisioning contact points
contactPoints:
  - orgId: 1
    name: slack-infrastructure
    receivers:
      - uid: slack-infra
        type: slack
        settings:
          url: "${SLACK_WEBHOOK_URL}"
          recipient: "#infrastructure-alerts"
          title: |
            {{ template "slack.default.title" . }}
          text: |
            {{ template "slack.default.text" . }}

# Provisioning notification policies
policies:
  - orgId: 1
    receiver: slack-infrastructure
    group_by:
      - alertname
      - severity
    routes:
      - receiver: pagerduty-critical
        matchers:
          - severity = critical
        continue: true
      - receiver: slack-infrastructure
        matchers:
          - team = infrastructure

Examples

Alert rule expressions:

# Multi-condition alert
data:
  - refId: A
    expr: sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
  - refId: B
    expr: sum(rate(http_requests_total[5m])) by (service)
  - refId: C
    datasourceUid: __expr__
    model:
      expression: A / B * 100
      type: math
  - refId: D
    datasourceUid: __expr__
    model:
      expression: C
      conditions:
        - evaluator:
            type: gt
            params: [5]
      type: threshold
condition: D

Contact point configurations:

# Email contact point
- type: email
  settings:
    addresses: oncall@example.com
    singleEmail: true

# PagerDuty contact point
- type: pagerduty
  settings:
    integrationKey: "${PAGERDUTY_KEY}"
    severity: critical
    class: infrastructure
    component: grafana

# Webhook contact point
- type: webhook
  settings:
    url: https://api.example.com/alerts
    httpMethod: POST
    username: grafana
    password: "${WEBHOOK_PASSWORD}"

Creating alerts via UI:

  1. Navigate to Alerting > Alert rules
  2. Click "New alert rule"
  3. Define query and conditions
  4. Set evaluation behaviour (folder, group, interval)
  5. Add labels and annotations
  6. Configure notification policy or contact point

Common Visualisation Types

Grafana provides numerous panel types for different data visualisation needs.

Key Concepts

Visualisation Best For Data Type
Time series Metrics over time Time series
Stat Single values, KPIs Scalar
Gauge Progress/percentage Scalar
Bar chart Comparisons Categorical
Table Detailed data Tabular
Heatmap Distribution over time Histogram
Logs Log entries Log streams
Node graph Relationships Graph data
Geomap Geographical data Coordinates

Common Patterns

Time series panel:

{
  "type": "timeseries",
  "fieldConfig": {
    "defaults": {
      "custom": {
        "drawStyle": "line",
        "lineInterpolation": "smooth",
        "fillOpacity": 10,
        "gradientMode": "scheme",
        "showPoints": "auto",
        "pointSize": 5,
        "stacking": { "mode": "none" },
        "axisPlacement": "auto",
        "spanNulls": true
      },
      "color": { "mode": "palette-classic" },
      "unit": "bytes",
      "decimals": 2,
      "min": 0
    }
  },
  "options": {
    "legend": {
      "displayMode": "table",
      "placement": "bottom",
      "calcs": ["mean", "max", "last"]
    },
    "tooltip": {
      "mode": "multi",
      "sort": "desc"
    }
  }
}

Stat panel:

{
  "type": "stat",
  "fieldConfig": {
    "defaults": {
      "unit": "percent",
      "thresholds": {
        "mode": "absolute",
        "steps": [
          { "value": null, "color": "green" },
          { "value": 80, "color": "yellow" },
          { "value": 90, "color": "red" }
        ]
      },
      "mappings": [
        {
          "type": "range",
          "options": {
            "from": 0,
            "to": 50,
            "result": { "text": "Low", "color": "green" }
          }
        }
      ]
    }
  },
  "options": {
    "reduceOptions": {
      "calcs": ["lastNotNull"],
      "fields": "",
      "values": false
    },
    "orientation": "horizontal",
    "textMode": "auto",
    "colorMode": "background",
    "graphMode": "area",
    "justifyMode": "center"
  }
}

Examples

Table panel with transformations:

{
  "type": "table",
  "transformations": [
    {
      "id": "organize",
      "options": {
        "excludeByName": { "Time": true },
        "renameByName": {
          "instance": "Host",
          "Value": "CPU %"
        }
      }
    },
    {
      "id": "sortBy",
      "options": {
        "fields": {},
        "sort": [{ "field": "CPU %", "desc": true }]
      }
    }
  ],
  "fieldConfig": {
    "overrides": [
      {
        "matcher": { "id": "byName", "options": "CPU %" },
        "properties": [
          { "id": "custom.cellOptions", "value": { "type": "color-background" } },
          { "id": "thresholds", "value": {
            "mode": "absolute",
            "steps": [
              { "value": null, "color": "green" },
              { "value": 70, "color": "yellow" },
              { "value": 90, "color": "red" }
            ]
          }}
        ]
      }
    ]
  }
}

Heatmap panel:

{
  "type": "heatmap",
  "options": {
    "calculate": false,
    "yAxis": {
      "axisPlacement": "left",
      "unit": "s"
    },
    "cellGap": 1,
    "color": {
      "scheme": "Spectral",
      "mode": "scheme",
      "fill": "dark-orange",
      "scale": "exponential",
      "exponent": 0.5
    },
    "tooltip": {
      "show": true,
      "yHistogram": true
    },
    "legend": { "show": true }
  }
}

Bar gauge panel:

{
  "type": "bargauge",
  "options": {
    "reduceOptions": {
      "calcs": ["lastNotNull"]
    },
    "orientation": "horizontal",
    "displayMode": "gradient",
    "showUnfilled": true,
    "minVizWidth": 0,
    "minVizHeight": 10
  }
}

Dashboard Organisation and Folders

Organise dashboards using folders, tags, and proper naming conventions for maintainability.

Key Concepts

Concept Description
Folders Container for grouping related dashboards
Tags Searchable labels applied to dashboards
Permissions Role-based access at folder and dashboard level
Starring Bookmark frequently used dashboards
Playlists Rotating display of dashboards

Common Patterns

# Folder provisioning
# /etc/grafana/provisioning/dashboards/default.yaml
apiVersion: 1

providers:
  - name: 'infrastructure'
    orgId: 1
    folder: 'Infrastructure'
    folderUid: 'infrastructure'
    type: file
    disableDeletion: false
    updateIntervalSeconds: 30
    allowUiUpdates: true
    options:
      path: /var/lib/grafana/dashboards/infrastructure

  - name: 'applications'
    orgId: 1
    folder: 'Applications'
    folderUid: 'applications'
    type: file
    options:
      path: /var/lib/grafana/dashboards/applications

  - name: 'business'
    orgId: 1
    folder: 'Business Metrics'
    folderUid: 'business'
    type: file
    options:
      path: /var/lib/grafana/dashboards/business

Recommended folder structure:

dashboards/
├── infrastructure/
│   ├── node-overview.json
│   ├── kubernetes-cluster.json
│   └── network-monitoring.json
├── applications/
│   ├── api-service.json
│   ├── web-frontend.json
│   └── worker-service.json
├── databases/
│   ├── postgresql.json
│   ├── redis.json
│   └── elasticsearch.json
├── business/
│   ├── revenue-metrics.json
│   └── user-engagement.json
└── sre/
    ├── slo-dashboard.json
    └── incident-response.json

Examples

# Create folder via API
curl -X POST \
  -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{"title": "Production"}' \
  http://localhost:3000/api/folders

# List folders
curl -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  http://localhost:3000/api/folders

# Set folder permissions
curl -X POST \
  -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "items": [
      {"role": "Viewer", "permission": 1},
      {"role": "Editor", "permission": 2},
      {"teamId": 1, "permission": 4}
    ]
  }' \
  http://localhost:3000/api/folders/abc123/permissions

# Search dashboards by tag
curl -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  "http://localhost:3000/api/search?tag=production&type=dash-db"

# Create playlist
curl -X POST \
  -H "Authorization: Bearer ${GRAFANA_SA_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "NOC Display",
    "interval": "5m",
    "items": [
      {"type": "dashboard_by_uid", "value": "abc123"},
      {"type": "dashboard_by_uid", "value": "def456"}
    ]
  }' \
  http://localhost:3000/api/playlists

Tagging conventions:

  • Environment: production, staging, development
  • Team: platform, backend, frontend, sre
  • Service type: api, database, cache, queue
  • Alert level: critical, warning, info

Provisioning Dashboards

Automate dashboard deployment using Grafana's provisioning system for version control and consistency.

Key Concepts

Provisioning TypesProvisioning FlowGit RepositoryCI/CD PipelineDeploy to Grafana/etc/grafana/provisioningDashboard ProviderGrafana DatabaseDashboardsData SourcesAlert RulesContact PointsNotificationPoliciesProvisioning TypesProvisioning FlowGit RepositoryCI/CD PipelineDeploy to Grafana/etc/grafana/provisioningDashboard ProviderGrafana DatabaseDashboardsData SourcesAlert RulesContact PointsNotificationPolicies

Provisioning directory structure:

/etc/grafana/provisioning/
├── dashboards/
│   ├── default.yaml       # Provider configuration
│   └── dashboards/        # Dashboard JSON files
├── datasources/
│   └── datasources.yaml   # Data source configurations
├── alerting/
│   ├── alerts.yaml        # Alert rules
│   └── contactpoints.yaml # Contact points
├── notifiers/             # Legacy alerting (deprecated)
└── plugins/
    └── plugins.yaml       # Plugin installations

Common Patterns

# Dashboard provider configuration
# /etc/grafana/provisioning/dashboards/default.yaml
apiVersion: 1

providers:
  - name: 'default'
    orgId: 1
    folder: ''
    folderUid: ''
    type: file
    disableDeletion: false
    updateIntervalSeconds: 10
    allowUiUpdates: false
    options:
      path: /var/lib/grafana/dashboards
      foldersFromFilesStructure: true
# Data source provisioning
# /etc/grafana/provisioning/datasources/datasources.yaml
apiVersion: 1

deleteDatasources:
  - name: OldPrometheus
    orgId: 1

datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    orgId: 1
    url: http://prometheus:9090
    isDefault: true
    version: 1
    editable: false
    jsonData:
      httpMethod: POST
      manageAlerts: true
      prometheusType: Prometheus
      prometheusVersion: 2.40.0

Examples

Dockerfile with provisioning:

FROM grafana/grafana:12.0.0

# Copy provisioning configurations
COPY provisioning/ /etc/grafana/provisioning/

# Copy dashboards
COPY dashboards/ /var/lib/grafana/dashboards/

# Copy plugins
COPY plugins/ /var/lib/grafana/plugins/

# Set environment variables
ENV GF_SECURITY_ADMIN_PASSWORD=admin \
    GF_USERS_ALLOW_SIGN_UP=false \
    GF_AUTH_ANONYMOUS_ENABLED=true \
    GF_AUTH_ANONYMOUS_ORG_ROLE=Viewer

Kubernetes ConfigMap for provisioning:

apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-dashboards-config
  namespace: monitoring
data:
  dashboards.yaml: |
    apiVersion: 1
    providers:
      - name: 'default'
        orgId: 1
        folder: ''
        type: file
        disableDeletion: false
        options:
          path: /var/lib/grafana/dashboards
---
apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-dashboards
  namespace: monitoring
  labels:
    grafana_dashboard: "1"
data:
  api-metrics.json: |
    {
      "dashboard": {
        "title": "API Metrics",
        ...
      }
    }

Helm values for Grafana provisioning:

# values.yaml for grafana helm chart
grafana:
  dashboardProviders:
    dashboardproviders.yaml:
      apiVersion: 1
      providers:
        - name: 'default'
          orgId: 1
          folder: ''
          type: file
          disableDeletion: false
          editable: true
          options:
            path: /var/lib/grafana/dashboards/default

  dashboards:
    default:
      node-exporter:
        gnetId: 1860
        revision: 27
        datasource: Prometheus
      kubernetes-cluster:
        gnetId: 6417
        revision: 1
        datasource: Prometheus
      custom-dashboard:
        file: dashboards/custom.json

  datasources:
    datasources.yaml:
      apiVersion: 1
      datasources:
        - name: Prometheus
          type: prometheus
          url: http://prometheus-server
          isDefault: true

  alerting:
    rules.yaml:
      apiVersion: 1
      groups:
        - orgId: 1
          name: my-alerts
          folder: alerts
          interval: 1m
          rules: []

Grafana Terraform provider:

resource "grafana_folder" "production" {
  title = "Production"
}

resource "grafana_dashboard" "api_metrics" {
  folder      = grafana_folder.production.id
  config_json = file("dashboards/api-metrics.json")

  depends_on = [grafana_data_source.prometheus]
}

resource "grafana_data_source" "prometheus" {
  type = "prometheus"
  name = "Prometheus"
  url  = "http://prometheus:9090"

  json_data_encoded = jsonencode({
    httpMethod = "POST"
  })
}

Quick Reference

Grafana CLI Commands

Command Description
grafana-cli plugins list-remote List available plugins
grafana-cli plugins install <plugin> Install a plugin
grafana-cli plugins update-all Update all plugins
grafana-cli admin reset-admin-password <password> Reset admin password
grafana cli admin secrets-migration re-encrypt Re-encrypt secrets with the configured encryption (idempotent)

The old grafana-cli admin data-migration encrypt-datasource-passwords is a legacy pre-v7 migration (moves plaintext data source passwords into secure_json_data) and is not needed on modern installs.

API Endpoints

Endpoint Method Description
/api/dashboards/db POST Create/update dashboard
/api/dashboards/uid/:uid GET Get dashboard by UID
/api/datasources GET/POST List/create data sources
/api/folders GET/POST List/create folders
/api/search GET Search dashboards
/api/v1/provisioning/alert-rules GET List alert rules
/api/annotations GET/POST Manage annotations
/api/org/users GET List organisation users

Common Variable Queries (Prometheus)

Purpose Query
All namespaces label_values(kube_pod_info, namespace)
Pods in namespace label_values(kube_pod_info{namespace="$namespace"}, pod)
All instances label_values(up, instance)
Jobs label_values(up, job)
Metric names label_names()
Label values label_values(metric_name, label_name)

Panel Configuration Options

Option Description
unit Display unit (bytes, percent, reqps, etc.)
decimals Number of decimal places
min/max Axis boundaries
thresholds Colour-coded value ranges
mappings Value to text/colour mappings
overrides Field-specific configurations
transformations Data manipulation steps

Environment Variables

Variable Description Default
GF_SECURITY_ADMIN_USER Admin username admin
GF_SECURITY_ADMIN_PASSWORD Admin password admin
GF_SERVER_ROOT_URL Full public URL -
GF_DATABASE_TYPE Database type sqlite3
GF_AUTH_ANONYMOUS_ENABLED Allow anonymous access false
GF_INSTALL_PLUGINS Comma-separated plugin list -
GF_PATHS_PROVISIONING Provisioning path /etc/grafana/provisioning

Common Issues and Solutions

Issue Cause Solution
Data source test fails Network/auth issues Verify URL, credentials, and network connectivity. Check Grafana server logs for detailed errors
No data in panel Query returns empty Use Query Inspector to debug. Check time range matches data. Verify data source is correct
Variables not loading Query syntax error Check variable query syntax. Ensure data source is selected. Look for errors in variable settings
Dashboard not updating Browser cache Hard refresh (Ctrl+Shift+R). Check auto-refresh interval. Verify data source is receiving data
Alerts not firing Condition never met Check alert rule evaluation in Alert Rules page. Verify expression in explore view. Check "for" duration
Provisioned dashboard shows "cannot save" Read-only provisioning Set allowUiUpdates: true in provider config, or edit source files directly
High memory usage Too many panels/queries Reduce panel count. Use recording rules for expensive queries. Implement pagination
Slow dashboard loading Complex queries Pre-aggregate with recording rules. Reduce time range. Use query caching
Permission denied RBAC restrictions Check user role and team permissions. Verify folder permissions
Plugin not found Not installed Install via CLI: grafana-cli plugins install <plugin-id>
Dashboard UID conflict Duplicate UIDs Generate new UID or remove conflicting dashboard before import
Template variable recursion Circular dependency Review variable dependencies. Ensure no variable references itself indirectly

Debugging Tips

# Check Grafana logs
journalctl -u grafana-server -f
# Or Docker logs
docker logs grafana -f

# Test data source connectivity
curl -v http://prometheus:9090/api/v1/query?query=up

# Check provisioning status
curl -H "Authorization: Bearer ${SA_TOKEN}" \
  http://localhost:3000/api/admin/provisioning/dashboards/reload

# Export dashboard for debugging
curl -H "Authorization: Bearer ${SA_TOKEN}" \
  http://localhost:3000/api/dashboards/uid/abc123 | jq

# Check plugin status
curl -H "Authorization: Bearer ${SA_TOKEN}" \
  http://localhost:3000/api/plugins

# Verify alert rule status
curl -H "Authorization: Bearer ${SA_TOKEN}" \
  http://localhost:3000/api/v1/provisioning/alert-rules

# Database backup (SQLite)
sqlite3 /var/lib/grafana/grafana.db ".backup '/backup/grafana.db'"

Query Inspector usage:

  1. Open panel in edit mode
  2. Click "Query Inspector" tab
  3. View "Query" for sent request
  4. View "Data" for response
  5. Check "Stats" for timing information
  6. Export results as CSV for analysis

Related Topics

The following topics complement Grafana and would enhance your monitoring and observability capabilities:

  1. Prometheus - Time-series database that pairs with Grafana as the primary metrics data source, providing PromQL for powerful queries
  2. Loki - Log aggregation system from Grafana Labs that integrates seamlessly with Grafana for log exploration and correlation with metrics
  3. Alertmanager - Handles alert routing, grouping, and notification dispatch for Prometheus alerts, complementing Grafana's alerting
  4. OpenTelemetry - Vendor-neutral observability framework for collecting traces, metrics, and logs that can be visualised in Grafana
  5. Tempo - Grafana's distributed tracing backend that integrates with Grafana for trace visualisation and correlation
  6. Kubernetes - Container orchestration platform commonly monitored with Grafana dashboards using Prometheus metrics and Loki logs