Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Loki

Log aggregation system designed for efficiency and cost-effectiveness, using labels for indexing rather than full-text indexing.

Loki

Log aggregation system designed for efficiency and cost-effectiveness, using labels for indexing rather than full-text indexing.

Overview

Loki is a horizontally scalable, highly available, multi-tenant log aggregation system inspired by Prometheus. Unlike traditional log systems that index the full content of logs, Loki only indexes metadata (labels), making it significantly cheaper to operate and easier to scale. It integrates seamlessly with Grafana and uses LogQL, a query language similar to PromQL, for querying logs.

Log SourcesLoki ArchitecturelogspushApplicationsAlloy/AgentsDistributorIngesterObject StorageIndex StoreQuery FrontendQuerierGrafanaDockerKubernetesSyslogFilesLog SourcesLoki ArchitecturelogspushApplicationsAlloy/AgentsDistributorIngesterObject StorageIndex StoreQuery FrontendQuerierGrafanaDockerKubernetesSyslogFiles

LogQL Query Language

LogQL is Loki's query language, combining log stream selection with filtering and aggregation capabilities.

Key Concepts

Query PipelineStream SelectorLine FiltersParserLabel FiltersMetric QueriesQuery PipelineStream SelectorLine FiltersParserLabel FiltersMetric Queries

Query Types:

  • Log queries: Return log lines matching criteria
  • Metric queries: Calculate values from log content over time

Stream Selectors use labels to filter log streams, similar to Prometheus label selectors.

Stream Selectors

# Exact match
{job="nginx"}

# Regex match
{job=~"nginx|apache"}

# Not equal
{job!="debug"}

# Negative regex match
{job!~"test.*"}

# Multiple labels
{job="api", environment="production", level="error"}

# Combine matchers
{namespace=~"prod-.*", container!="sidecar"}

Line Filters

# Contains string (case-sensitive)
{job="nginx"} |= "error"

# Does not contain
{job="nginx"} != "debug"

# Regex match
{job="nginx"} |~ "status=[45][0-9]{2}"

# Negative regex match
{job="nginx"} !~ "health_check"

# Chain multiple filters (AND logic)
{job="api"} |= "error" != "timeout" |~ "user_id=[0-9]+"

# Case-insensitive match
{job="nginx"} |~ "(?i)error"

Parsers

# JSON parser - extracts all JSON fields as labels
{job="api"} | json

# Extract specific JSON fields
{job="api"} | json level, message, user_id

# Logfmt parser
{job="api"} | logfmt

# Regex parser with named groups
{job="nginx"} | regexp `(?P<ip>\d+\.\d+\.\d+\.\d+) - - \[(?P<timestamp>[^\]]+)\] "(?P<method>\w+) (?P<path>[^"]+)"`

# Pattern parser (simplified regex)
{job="nginx"} | pattern `<ip> - - [<timestamp>] "<method> <path> <_>" <status> <size>`

# Unpack parser for logs packed with stage.pack (Alloy) or Promtail pack stage
{job="api"} | unpack

# Line format - restructure log line
{job="api"} | json | line_format "{{.level}} - {{.message}}"

Label Filters

# Filter on extracted labels
{job="api"} | json | level="error"

# Numeric comparisons
{job="api"} | json | status >= 400

# Duration comparisons
{job="api"} | json | response_time > 500ms

# Byte comparisons
{job="api"} | json | size > 1kb

# Combine filters
{job="api"} | json | level="error" and status >= 500

# OR logic
{job="api"} | json | level="error" or level="warn"

# Regex on extracted labels
{job="api"} | json | message =~ ".*timeout.*"

Metric Queries

# Count logs per second over time
rate({job="nginx"}[5m])

# Count specific errors
rate({job="nginx"} |= "error" [1m])

# Sum by label
sum by (level) (rate({job="api"} | json [5m]))

# Count unique values
count_over_time({job="api"}[1h])

# Bytes rate
bytes_rate({job="nginx"}[5m])

# Bytes over time
bytes_over_time({job="nginx"}[1h])

# Quantile from extracted values
quantile_over_time(0.95, {job="api"} | json | unwrap response_time [5m])

# Average of extracted numeric field
avg_over_time({job="api"} | json | unwrap duration [5m]) by (endpoint)

# Absent over time (detect missing logs)
absent_over_time({job="critical-service"}[5m])

Advanced Aggregations

# Top 5 paths by request count
topk(5, sum by (path) (rate({job="nginx"} | pattern `<_> "<_> <path> <_>"` [5m])))

# Error rate percentage
sum(rate({job="api"} | json | level="error" [5m]))
  / sum(rate({job="api"} [5m])) * 100

# Aggregate and sort by multiple dimensions
sum by (service, level) (count_over_time({namespace="production"} | json [1h]))

# Compare with offset
sum(rate({job="api"} |= "error" [1h]))
  - sum(rate({job="api"} |= "error" [1h] offset 24h))

# Group left/right for joining
sum by (instance) (rate({job="api"} |= "error" [5m]))
  / on (instance) group_left
sum by (instance) (rate({job="api"} [5m]))

Examples

# Find all 5xx errors in nginx logs
{job="nginx"} |~ "HTTP/[0-9.]+ [5][0-9]{2}"

# Parse structured API logs and filter
{job="api", environment="production"}
  | json
  | level="error"
  | line_format "{{.timestamp}} [{{.level}}] {{.message}}"

# Calculate p99 response time from JSON logs
quantile_over_time(0.99,
  {job="api"}
    | json
    | unwrap response_ms
    | __error__=""
  [5m]
) by (endpoint)

# Count errors by service and error type
sum by (service, error_type) (
  count_over_time(
    {namespace="production"}
      | json
      | level="error"
    [1h]
  )
)

# Find slow requests
{job="api"}
  | json
  | response_time > 1s
  | line_format "{{.method}} {{.path}} took {{.response_time}}"

Label Strategies

Effective label design is crucial for Loki performance and query efficiency.

Key Concepts

Bad LabelsGood LabelsStatic valuesLow cardinalityUseful for filteringDynamic valuesHigh cardinalityUnique identifiersGood PerformancePoor PerformanceBad LabelsGood LabelsStatic valuesLow cardinalityUseful for filteringDynamic valuesHigh cardinalityUnique identifiersGood PerformancePoor Performance

Cardinality Guidelines:

  • Keep total unique label combinations under 10,000 per tenant
  • Each unique label set creates a new stream
  • High cardinality causes increased memory usage and query times

Common Label Patterns

# Recommended labels
labels:
  job: nginx                    # Application name
  environment: production       # Environment
  namespace: web-services       # Kubernetes namespace
  cluster: eu-west-1           # Cluster/region
  level: error                 # Log level (if consistent)

# Avoid these as labels
labels:
  request_id: "abc-123"        # High cardinality - unique per request
  user_id: "12345"             # High cardinality - unique per user
  timestamp: "2024-01-01T..."  # Already stored in log entry
  ip_address: "10.0.0.1"       # High cardinality - many unique IPs
  trace_id: "xyz-789"          # High cardinality - unique per trace

Label Design Patterns

# Pattern 1: Environment-based hierarchy
labels:
  environment: production
  region: eu-west-1
  cluster: main
  namespace: api
  service: user-service
  pod: user-service-abc123     # Only if truly needed

# Pattern 2: Team ownership
labels:
  team: platform
  service: auth
  component: api

# Pattern 3: Log categorisation
labels:
  job: application
  level: info                  # Only if log level is static per stream
  format: json                 # Helps choose parser

# Pattern 4: Kubernetes standard labels
labels:
  namespace: default
  deployment: my-app
  container: main
  pod: my-app-xyz              # Consider if this adds value

Dynamic Labelling Strategies

// Alloy: extract level as label using loki.process (use cautiously — high cardinality risk)
loki.process "api_labels" {
  forward_to = [loki.write.default.receiver]

  stage.json {
    expressions = { level = "level" }
  }

  stage.labels {
    values = { level = "" }
  }
}

loki.source.file "api" {
  targets = [{
    __path__ = "/var/log/api/*.log",
    job      = "api",
  }]
  forward_to = [loki.process.api_labels.receiver]
}

// Drop high-cardinality labels after discovery
loki.process "nginx_drop" {
  forward_to = [loki.write.default.receiver]

  stage.label_drop {
    values = ["pod_template_hash", "controller_revision_hash"]
  }
}

Stream Optimisation

// Consolidate similar streams
// Instead of labels: {app="api", instance="pod-1"}, {app="api", instance="pod-2"} …
// Consider:          {app="api"}
// Use structured metadata or line content for per-instance identification.

// Alloy: keep essential labels only via discovery.relabel
discovery.relabel "k8s_essential" {
  targets = discovery.kubernetes.pods.targets

  rule {
    action        = "labelkeep"
    regex         = "(namespace|container|pod)"
  }

  rule {
    source_labels = ["__meta_kubernetes_namespace"]
    target_label  = "namespace"
  }
}

Grafana Alloy — Log Collection Agent

Grafana Alloy is the recommended log collection agent for Loki. It supersedes Promtail, which entered Long-Term Support on 13 February 2025 and reached End-of-Life on 2 March 2026. All new deployments should use Alloy.

Migration: alloy convert --source-format=promtail --output=config.alloy promtail.yaml converts an existing Promtail config to Alloy format automatically. You can also run Alloy directly against a Promtail config file with alloy run --config.format=promtail promtail.yaml while you migrate. See the official migration guide.

Legacy Promtail: if you still operate Promtail in a legacy environment, the pipeline-stage concepts (JSON, regex, timestamp, labels, multiline, drop, match, replace, pack, tenant) map directly to Alloy stage.* blocks inside loki.process. The syntax differs but the semantics are identical.

Key Concepts

Pipeline Stages (loki.process)Service DiscoveryTarget Files / K8sPodsloki.source.file /loki.source.kubernetesloki.process —Pipeline Stagesloki.write — Push toLokistage.json /stage.regexstage.timestampstage.labels /stage.static_labelsstage.drop /stage.matchstage.output /stage.replacePipeline Stages (loki.process)Service DiscoveryTarget Files / K8sPodsloki.source.file /loki.source.kubernetesloki.process —Pipeline Stagesloki.write — Push toLokistage.json /stage.regexstage.timestampstage.labels /stage.static_labelsstage.drop /stage.matchstage.output /stage.replace

Core Alloy components for log collection:

  • local.file_match: Discovers files on disk using glob patterns
  • loki.source.file: Tails discovered files and forwards log entries
  • loki.source.kubernetes: Tails Kubernetes pod logs via the API (no node privileges needed)
  • discovery.kubernetes: Discovers Kubernetes resources for targeting
  • discovery.relabel: Applies relabelling rules to discovered targets
  • loki.process: Applies processing stages (parse, transform, filter)
  • loki.write: Ships logs to a Loki endpoint

Basic File Collection

// config.alloy — minimal file collection

// Discover log files
local.file_match "system" {
  path_targets = [
    { __path__ = "/var/log/*.log", job = "varlogs" },
  ]
}

// Tail matched files and forward to Loki
loki.source.file "system" {
  targets    = local.file_match.system.targets
  forward_to = [loki.write.default.receiver]
}

// Multiple path patterns
local.file_match "applications" {
  path_targets = [
    { __path__ = "/var/log/apps/**/*.log", job = "apps" },
  ]
}

loki.source.file "applications" {
  targets    = local.file_match.applications.targets
  forward_to = [loki.write.default.receiver]
}

// Loki endpoint
loki.write "default" {
  endpoint {
    url       = "http://loki:3100/loki/api/v1/push"
    tenant_id = "default"
  }
  external_labels = {
    cluster = "production-eu",
  }
}

Processing Stages (loki.process)

// JSON API logs: parse, timestamp, label, output, redact

local.file_match "api" {
  path_targets = [{
    __path__     = "/var/log/api/*.log",
    job          = "api",
    environment  = "production",
  }]
}

loki.source.file "api" {
  targets    = local.file_match.api.targets
  forward_to = [loki.process.api.receiver]
}

loki.process "api" {
  forward_to = [loki.write.default.receiver]

  // Parse JSON log line
  stage.json {
    expressions = {
      level     = "level",
      message   = "message",
      ts        = "timestamp",
      duration  = "response_time",
    }
  }

  // Set timestamp from extracted field
  // Supported shorthand: RFC3339, RFC3339Nano, Unix, UnixMs, UnixUs, UnixNs
  stage.timestamp {
    source = "ts"
    format = "RFC3339Nano"
  }

  // Promote extracted fields to labels (low-cardinality only)
  stage.labels {
    values = { level = "" }
  }

  // Add static labels
  stage.static_labels {
    values = { environment = "production" }
  }

  // Rewrite the log line to just the message field
  stage.output {
    source = "message"
  }
}

// Nginx access log: regex parsing
local.file_match "nginx" {
  path_targets = [{
    __path__ = "/var/log/nginx/access.log",
    job      = "nginx",
  }]
}

loki.source.file "nginx" {
  targets    = local.file_match.nginx.targets
  forward_to = [loki.process.nginx.receiver]
}

loki.process "nginx" {
  forward_to = [loki.write.default.receiver]

  // Use backtick raw strings for regex — no double-escaping needed
  stage.regex {
    expression = `^(?P<ip>\S+) - (?P<user>\S+) \[(?P<ts>[^\]]+)\] "(?P<method>\S+) (?P<path>\S+) (?P<proto>[^"]+)" (?P<status>\d+) (?P<size>\d+)`
  }

  stage.timestamp {
    source = "ts"
    format = "02/Jan/2006:15:04:05 -0700"
  }

  // Derive status group (2xx, 4xx, 5xx) with template
  stage.template {
    source   = "status_group"
    template = `{{ .status | substr 0 1 }}xx`
  }

  stage.labels {
    values = { status_group = "" }
  }
}

Advanced Processing Stages

loki.process "advanced" {
  forward_to = [loki.write.default.receiver]

  // Multiline logs — collapse stack traces into a single entry
  stage.multiline {
    firstline     = `^\d{4}-\d{2}-\d{2}`
    max_wait_time = "3s"
    max_lines     = 128
  }

  // Drop health-check noise
  stage.match {
    selector            = `{job="nginx"} |~ ".*health_check.*"`
    action              = "drop"
    drop_counter_reason = "health_check"
  }

  // Derive metrics from log content
  stage.metrics {
    metric.counter {
      name        = "http_requests_total"
      description = "Total HTTP requests"
      prefix      = "alloy_custom_"
      source      = "status"
      action      = "inc"
    }

    metric.histogram {
      name        = "request_duration_seconds"
      description = "Request duration"
      prefix      = "alloy_custom_"
      source      = "duration"
      buckets     = [0.01, 0.05, 0.1, 0.5, 1, 5]
    }
  }

  // Redact sensitive data
  stage.replace {
    expression = `(password=)[^&\s]+`
    replace    = "${1}***REDACTED***"
  }

  // Pack high-cardinality fields into the log line body instead of labels
  stage.pack {
    labels            = ["level", "service"]
    ingest_timestamp  = true
  }

  // Set tenant from an extracted label
  stage.tenant {
    source = "team"
  }
}

Kubernetes Log Collection

// Pod logs via the Kubernetes API — no DaemonSet node privileges required.
// Use loki.source.kubernetes for API-based tailing,
// or loki.source.file (DaemonSet) for file-based tailing from /var/log/pods.

// --- Option A: API-based (recommended for most clusters) ---

discovery.kubernetes "pods" {
  role = "pod"

  // When running as a DaemonSet, restrict to pods on this node
  selectors {
    role  = "pod"
    field = "spec.nodeName=" + coalesce(sys.env("HOSTNAME"), constants.hostname)
  }
}

discovery.relabel "pod_logs" {
  targets = discovery.kubernetes.pods.targets

  // Only scrape pods with the opt-in annotation
  rule {
    source_labels = ["__meta_kubernetes_pod_annotation_alloy_io_scrape"]
    action        = "keep"
    regex         = "true"
  }

  rule {
    source_labels = ["__meta_kubernetes_namespace"]
    target_label  = "namespace"
  }

  rule {
    source_labels = ["__meta_kubernetes_pod_name"]
    target_label  = "pod"
  }

  rule {
    source_labels = ["__meta_kubernetes_pod_container_name"]
    target_label  = "container"
  }

  rule {
    source_labels = ["__meta_kubernetes_pod_label_app_kubernetes_io_name"]
    target_label  = "app"
  }

  rule {
    source_labels = ["__meta_kubernetes_namespace", "__meta_kubernetes_pod_container_name"]
    separator     = "/"
    replacement   = "$1"
    target_label  = "job"
  }

  // Drop verbose Kubernetes meta-labels from the final label set
  rule {
    action = "labeldrop"
    regex  = "__meta_kubernetes_pod_label_(.+)"
  }
}

loki.source.kubernetes "pod_logs" {
  targets    = discovery.relabel.pod_logs.output
  forward_to = [loki.process.pod_logs.receiver]
}

loki.process "pod_logs" {
  forward_to = [loki.write.default.receiver]

  // CRI format is standard for containerd / CRI-O runtimes
  stage.cri {}

  // Add cluster-level static label
  stage.static_labels {
    values = { cluster = "production-eu" }
  }
}

// --- Option B: file-based DaemonSet (requires host /var/log/pods mount) ---

discovery.relabel "pod_file_logs" {
  targets = discovery.kubernetes.pods.targets

  rule {
    source_labels = ["__meta_kubernetes_pod_uid", "__meta_kubernetes_pod_container_name"]
    separator     = "/"
    replacement   = "/var/log/pods/*$1/*.log"
    target_label  = "__path__"
  }

  rule {
    source_labels = ["__meta_kubernetes_namespace"]
    target_label  = "namespace"
  }

  rule {
    source_labels = ["__meta_kubernetes_pod_name"]
    target_label  = "pod"
  }

  rule {
    source_labels = ["__meta_kubernetes_pod_container_name"]
    target_label  = "container"
  }
}

loki.source.file "pod_file_logs" {
  targets    = discovery.relabel.pod_file_logs.output
  forward_to = [loki.process.pod_logs.receiver]
}

Complete Production Example

// config.alloy — production application logs

// --- Loki endpoint ---
loki.write "default" {
  endpoint {
    url = "http://loki-gateway:3100/loki/api/v1/push"

    // For Grafana Cloud / authenticated endpoints:
    // basic_auth {
    //   username = "<USERNAME>"
    //   password = "<API_KEY>"
    // }
  }

  external_labels = {
    cluster = "production-eu",
  }
}

// --- Application logs with multiline + JSON pipeline ---
local.file_match "application" {
  path_targets = [{
    __path__     = "/var/log/app/*.log",
    job          = "application",
    environment  = "production",
  }]
}

loki.source.file "application" {
  targets    = local.file_match.application.targets
  forward_to = [loki.process.application.receiver]
}

loki.process "application" {
  forward_to = [loki.write.default.receiver]

  // Collapse multiline entries (ISO 8601 timestamp starts a new entry)
  stage.multiline {
    firstline     = `^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}`
    max_wait_time = "3s"
  }

  stage.json {
    expressions = {
      level    = "level",
      service  = "service",
      trace_id = "trace_id",
      ts       = "@timestamp",
    }
  }

  stage.timestamp {
    source = "ts"
    format = "RFC3339Nano"
  }

  stage.labels {
    values = {
      level   = "",
      service = "",
    }
  }

  // trace_id is high-cardinality — store as structured metadata, not a label
  stage.structured_metadata {
    values = { trace_id = "" }
  }

  // Drop debug logs in production
  stage.match {
    selector            = `{level="debug"}`
    action              = "drop"
    drop_counter_reason = "debug_logs"
  }
}

// --- Kubernetes cluster events ---
loki.source.kubernetes_events "cluster_events" {
  job_name   = "integrations/kubernetes/eventhandler"
  log_format = "logfmt"
  forward_to = [loki.process.cluster_events.receiver]
}

loki.process "cluster_events" {
  forward_to = [loki.write.default.receiver]

  stage.static_labels {
    values = { cluster = "production-eu" }
  }
}

Integration with Grafana

Grafana provides the primary visualisation interface for Loki, enabling log exploration, dashboards, and alerting.

Key Concepts

StorageLokiGrafanaUserStorageLokiGrafanaUserQuery logsLogQL queryFetch chunksLog dataProcess & filterResultsVisualisationStorageLokiGrafanaUserStorageLokiGrafanaUserQuery logsLogQL queryFetch chunksLog dataProcess & filterResultsVisualisation

Integration Features:

  • Explore view for ad-hoc queries
  • Dashboard panels for visualisation
  • Alerting on log patterns
  • Log context and correlations

Data Source Configuration

# Grafana datasource provisioning
apiVersion: 1

datasources:
  - name: Loki
    type: loki
    access: proxy
    url: http://loki:3100
    isDefault: false
    jsonData:
      maxLines: 1000
      derivedFields:
        # Link to traces
        - name: TraceID
          matcherRegex: '"trace_id":"(\w+)"'
          url: '$${__value.raw}'
          datasourceUid: tempo
          urlDisplayLabel: View Trace
        # Link to external system
        - name: RequestID
          matcherRegex: 'request_id=(\w+)'
          url: 'https://logs.example.com/request/$${__value.raw}'
          urlDisplayLabel: View in External System
    secureJsonData:
      # For multi-tenant setups
      httpHeaderValue1: tenant-id

Dashboard Panels

{
  "panels": [
    {
      "title": "Error Logs",
      "type": "logs",
      "datasource": "Loki",
      "targets": [
        {
          "expr": "{job=\"api\"} |= \"error\" | json",
          "refId": "A"
        }
      ],
      "options": {
        "showTime": true,
        "showLabels": true,
        "wrapLogMessage": true,
        "prettifyLogMessage": true,
        "enableLogDetails": true,
        "dedupStrategy": "none",
        "sortOrder": "Descending"
      }
    },
    {
      "title": "Log Volume",
      "type": "timeseries",
      "datasource": "Loki",
      "targets": [
        {
          "expr": "sum by (level) (rate({job=\"api\"} | json [5m]))",
          "legendFormat": "{{level}}",
          "refId": "A"
        }
      ]
    },
    {
      "title": "Error Rate",
      "type": "stat",
      "datasource": "Loki",
      "targets": [
        {
          "expr": "sum(rate({job=\"api\"} |= \"error\" [5m])) / sum(rate({job=\"api\"} [5m])) * 100",
          "refId": "A"
        }
      ],
      "options": {
        "colorMode": "value",
        "graphMode": "none"
      },
      "fieldConfig": {
        "defaults": {
          "unit": "percent",
          "thresholds": {
            "mode": "absolute",
            "steps": [
              {"color": "green", "value": null},
              {"color": "yellow", "value": 1},
              {"color": "red", "value": 5}
            ]
          }
        }
      }
    }
  ]
}

Alerting Rules

# Loki ruler configuration
groups:
  - name: application-alerts
    interval: 1m
    rules:
      # Alert on error rate
      - alert: HighErrorRate
        expr: |
          sum(rate({job="api"} |= "error" [5m]))
          / sum(rate({job="api"} [5m])) > 0.05
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High error rate in API logs"
          description: "Error rate is {{ $value | humanizePercentage }}"

      # Alert on specific error pattern
      - alert: DatabaseConnectionError
        expr: |
          count_over_time({job="api"} |= "database connection failed" [5m]) > 5
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Database connection errors detected"
          description: "{{ $value }} connection errors in last 5 minutes"

      # Alert on missing logs
      - alert: NoLogsReceived
        expr: |
          absent_over_time({job="critical-service"}[5m])
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "No logs from critical service"
          description: "No logs received from critical-service for 5 minutes"

Correlating with Metrics and Traces

# Grafana data source correlations
datasources:
  - name: Loki
    type: loki
    jsonData:
      derivedFields:
        # Trace correlation
        - name: traceID
          matcherRegex: '"traceID":"(\w+)"'
          url: '$${__value.raw}'
          datasourceUid: tempo
        # Link to metrics
        - name: service
          matcherRegex: '"service":"(\w+)"'
          url: '/explore?left={"datasource":"prometheus","queries":[{"expr":"up{service=\"$${__value.raw}\"}"}]}'
          urlDisplayLabel: View Metrics

Explore View Tips

# Use query builder for complex queries
# Start with stream selector
{job="api", namespace="production"}

# Add filters progressively
{job="api"} |= "error"

# Parse and filter
{job="api"} | json | level="error" | status >= 500

# Add time range in URL
/explore?left={"range":{"from":"now-1h","to":"now"},"queries":[...]}

# Split view for comparison
# Use right panel for different query or time range

Log Aggregation Patterns

Design patterns for collecting, routing, and managing logs at scale.

Key Concepts

StorageRoutingCollection PatternsDirect PushSidecarDaemonSetAggregatorBy LabelBy ContentBy TenantHotWarmColdStorageRoutingCollection PatternsDirect PushSidecarDaemonSetAggregatorBy LabelBy ContentBy TenantHotWarmCold

Collection Architectures

# Pattern 1: DaemonSet (Kubernetes) — Alloy per node, file-based log access
# Install via the official Grafana Alloy Helm chart:
#   helm upgrade --install alloy grafana/alloy -f alloy-values.yaml
# The chart mounts /var/log and /var/lib/docker/containers automatically.
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: alloy
spec:
  template:
    spec:
      containers:
        - name: alloy
          image: grafana/alloy:latest
          args:
            - run
            - /etc/alloy/config.alloy
          volumeMounts:
            - name: config
              mountPath: /etc/alloy
            - name: varlog
              mountPath: /var/log
              readOnly: true
            - name: containers
              mountPath: /var/lib/docker/containers
              readOnly: true
      volumes:
        - name: config
          configMap:
            name: alloy-config
        - name: varlog
          hostPath:
            path: /var/log
        - name: containers
          hostPath:
            path: /var/lib/docker/containers

---
# Pattern 2: Sidecar (per-pod) — Alloy sidecar for apps with custom log paths
apiVersion: v1
kind: Pod
metadata:
  name: app-with-alloy-sidecar
spec:
  containers:
    - name: app
      image: myapp
      volumeMounts:
        - name: logs
          mountPath: /var/log/app
    - name: alloy
      image: grafana/alloy:latest
      args:
        - run
        - /etc/alloy/config.alloy
      volumeMounts:
        - name: logs
          mountPath: /var/log/app
          readOnly: true
        - name: config
          mountPath: /etc/alloy
  volumes:
    - name: logs
      emptyDir: {}
    - name: config
      configMap:
        name: alloy-sidecar-config

---
# Pattern 3: Aggregator — Alloy fan-out to multiple Loki clusters
# In config.alloy, define multiple loki.write endpoints:
#
# loki.write "primary" {
#   endpoint { url = "http://loki-primary:3100/loki/api/v1/push" }
# }
# loki.write "secondary" {
#   endpoint { url = "http://loki-secondary:3100/loki/api/v1/push" }
#   external_labels = { replica = "secondary" }
# }
#
# Then forward_to = [loki.write.primary.receiver, loki.write.secondary.receiver]

Multi-Tenant Patterns

// Alloy: static tenant ID on the loki.write endpoint
loki.write "team_a" {
  endpoint {
    url       = "http://loki:3100/loki/api/v1/push"
    tenant_id = "team-a"
  }
  external_labels = { tenant = "team-a" }
}

// Dynamic tenant ID from an extracted label
loki.process "multi_tenant" {
  forward_to = [loki.write.team_a.receiver]

  stage.json {
    expressions = { team = "team" }
  }

  // Sets the X-Scope-OrgID header from the extracted 'team' value
  stage.tenant {
    source = "team"
  }
}
# Loki server: multi-tenant configuration
auth_enabled: true
limits_config:
  ingestion_rate_mb: 10
  ingestion_burst_size_mb: 20
  max_streams_per_user: 10000
  max_global_streams_per_user: 50000

Log Routing and Filtering

// Alloy: route, filter, and sample logs based on content

loki.process "router" {
  // Fan-out: send to both primary and secondary clusters simultaneously
  forward_to = [
    loki.write.primary.receiver,
    loki.write.secondary.receiver,
  ]

  stage.json {
    expressions = {
      level   = "level",
      service = "service",
    }
  }

  // Route: promote level and service as labels for stream selection
  stage.labels {
    values = {
      level   = "",
      service = "",
    }
  }

  // Drop debug logs in production
  stage.match {
    selector            = `{level="debug"}`
    action              = "drop"
    drop_counter_reason = "debug_dropped"
  }

  // Sample verbose info logs — keep 10%
  stage.match {
    selector = `{level="info"}`

    stage.sampling {
      rate                = 0.1
      drop_counter_reason = "info_sampled"
    }
  }
}

loki.write "primary" {
  endpoint {
    url = "http://loki-primary:3100/loki/api/v1/push"
  }
}

loki.write "secondary" {
  endpoint {
    url = "http://loki-secondary:3100/loki/api/v1/push"
  }
  external_labels = { replica = "secondary" }
}

Retention and Lifecycle

# Loki retention configuration
# Loki 3.0 defaults to the TSDB index and schema v13; boltdb-shipper
# and schema v11 are legacy and not recommended for new deployments.
schema_config:
  configs:
    - from: 2024-01-01
      store: tsdb
      object_store: s3
      schema: v13
      index:
        prefix: loki_index_
        period: 24h

compactor:
  working_directory: /data/loki/compactor
  # `shared_store` was removed in Loki 3.0. The compactor now uses the
  # object store from storage_config; set delete_request_store (required
  # when retention is enabled) to your object store.
  delete_request_store: s3
  compaction_interval: 10m
  retention_enabled: true
  retention_delete_delay: 2h
  retention_delete_worker_count: 150

limits_config:
  retention_period: 720h  # 30 days default

# Per-stream retention
overrides:
  tenant-a:
    retention_period: 2160h  # 90 days for tenant-a
  tenant-b:
    retention_period: 168h   # 7 days for tenant-b

Performance Optimisation

Strategies for improving Loki query performance, ingestion throughput, and resource efficiency.

Key Concepts

Optimisation AreasPerformance FactorsQuery ScopePerformanceLabel DesignChunk ConfigIngestionStorageQueryingOptimisation AreasPerformance FactorsQuery ScopePerformanceLabel DesignChunk ConfigIngestionStorageQuerying

Performance Principles:

  • Minimise label cardinality
  • Use appropriate chunk sizing
  • Cache effectively
  • Parallelise queries

Query Optimisation

# Bad: No stream selector
|= "error"

# Good: Specific stream selector
{job="api", namespace="production"} |= "error"

# Bad: Wide time range
{job="api"}  # defaults to last hour or more

# Good: Narrow time range
{job="api"} # with explicit short time range in query

# Bad: Inefficient filter order
{job="api"} |~ "user_id=\d+" |= "error"

# Good: Most selective filter first
{job="api"} |= "error" |~ "user_id=\d+"

# Bad: Unnecessary parsing
{job="api"} | json | line_format "{{.message}}"

# Good: Only parse if needed for filtering
{job="api"} |= "error" | json | level="error"

# Use bloom filters (if enabled)
{job="api"} |= "error"  # Exact match uses bloom filter

Ingestion Optimisation

# Loki ingester configuration
ingester:
  chunk_idle_period: 30m       # Flush inactive chunks
  chunk_block_size: 262144     # 256KB uncompressed
  chunk_target_size: 1572864   # 1.5MB compressed target
  chunk_retain_period: 30s     # Keep chunks in memory after flush
  max_transfer_retries: 0      # Disable transfers in StatefulSet

  wal:
    enabled: true
    dir: /data/loki/wal
    flush_on_shutdown: true
    replay_memory_ceiling: 4GB

# Distributor configuration
distributor:
  ring:
    kvstore:
      store: memberlist

# Limits for ingestion
limits_config:
  ingestion_rate_mb: 20
  ingestion_burst_size_mb: 30
  per_stream_rate_limit: 5MB
  per_stream_rate_limit_burst: 15MB
  max_line_size: 256kb
  max_entries_limit_per_query: 5000

Storage Optimisation

# Chunk storage configuration
# TSDB is the recommended index store from Loki 2.8+ (default in 3.0).
# The `shared_store` field was removed in Loki 3.0; the shipper now
# derives the object store from the storage_config backend below.
storage_config:
  tsdb_shipper:
    active_index_directory: /data/loki/tsdb-index
    cache_location: /data/loki/tsdb-cache
    cache_ttl: 24h

  aws:
    s3: s3://region/bucket-name
    s3forcepathstyle: false

  # Filesystem for development
  filesystem:
    directory: /data/loki/chunks

# Caching
chunk_store_config:
  max_look_back_period: 0s
  chunk_cache_config:
    embedded_cache:
      enabled: true
      max_size_mb: 500
      ttl: 1h

query_range:
  results_cache:
    cache:
      embedded_cache:
        enabled: true
        max_size_mb: 100

# Index caching
index_queries_cache_config:
  embedded_cache:
    enabled: true
    max_size_mb: 100

Query Frontend Optimisation

# Query frontend configuration
frontend:
  max_outstanding_per_tenant: 2048
  compress_responses: true
  log_queries_longer_than: 5s

query_scheduler:
  max_outstanding_requests_per_tenant: 2048

querier:
  max_concurrent: 10
  query_ingesters_within: 3h  # Only query ingesters for recent data

query_range:
  align_queries_with_step: true
  max_retries: 5
  cache_results: true
  parallelise_shardable_queries: true

# Split queries by time interval
limits_config:
  split_queries_by_interval: 30m
  max_query_parallelism: 32
  max_query_length: 721h  # 30 days
  max_query_series: 500

Resource Tuning

# Kubernetes resource configuration
resources:
  # Ingester - memory intensive
  ingester:
    requests:
      cpu: "1"
      memory: 4Gi
    limits:
      cpu: "2"
      memory: 8Gi

  # Querier - CPU intensive
  querier:
    requests:
      cpu: "2"
      memory: 2Gi
    limits:
      cpu: "4"
      memory: 4Gi

  # Query Frontend - moderate resources
  query-frontend:
    requests:
      cpu: "500m"
      memory: 1Gi
    limits:
      cpu: "1"
      memory: 2Gi

# JVM tuning for components using BoltDB
# Environment variables
GOGC: "80"  # Tune garbage collection
GOMEMLIMIT: "6GiB"  # Soft memory limit

Monitoring Loki Performance

# Ingestion rate
sum(rate(loki_distributor_bytes_received_total[5m]))

# Query latency
histogram_quantile(0.99, sum(rate(loki_request_duration_seconds_bucket{route=~"loki_api_v1_query.*"}[5m])) by (le))

# Chunk flush rate
sum(rate(loki_ingester_chunks_flushed_total[5m]))

# Cache hit rate
sum(rate(loki_cache_hits_total[5m])) / sum(rate(loki_cache_fetched_keys_total[5m]))

# Stream count per tenant
loki_ingester_streams_created_total - loki_ingester_streams_removed_total

# Query queue length
loki_query_scheduler_queue_length

Quick Reference

LogQL Syntax

Element Syntax Example
Stream selector {label="value"} {job="nginx"}
Exact match = {env="prod"}
Regex match =~ {job=~"api|web"}
Not equal != {env!="dev"}
Contains |= {job="api"} |= "error"
Not contains != {job="api"} != "debug"
Regex line |~ {job="api"} |~ "status=[45].."
JSON parser | json {job="api"} | json
Label filter | field="value" {job="api"} | json | level="error"
Rate rate({...}[5m]) rate({job="api"}[5m])
Count count_over_time({...}[1h]) count_over_time({job="api"} |= "error"[1h])

Common LogQL Patterns

Pattern Query
Error count count_over_time({job="api"} |= "error" [1h])
Error rate rate({job="api"} |= "error" [5m])
Top errors by type topk(10, sum by (error) (count_over_time({job="api"} | json [1h])))
Log volume by level sum by (level) (rate({job="api"} | json [5m]))
P99 latency quantile_over_time(0.99, {job="api"} | json | unwrap duration [5m])
Missing logs absent_over_time({job="critical"}[5m])

Alloy Processing Stages (loki.process)

Stage block Purpose Example
stage.json Parse JSON stage.json { expressions = { level = "level" } }
stage.logfmt Parse logfmt stage.logfmt { mapping = { level = "" } }
stage.regex Parse with regex stage.regex { expression = \(?P<ip>\S+)` }`
stage.timestamp Set timestamp stage.timestamp { source = "ts" format = "RFC3339" }
stage.labels Add dynamic labels stage.labels { values = { level = "" } }
stage.static_labels Add static labels stage.static_labels { values = { env = "prod" } }
stage.output Rewrite log line stage.output { source = "message" }
stage.drop Drop lines stage.drop { expression = ".*health.*" }
stage.match Conditional stages stage.match { selector = \{job="api"}` action = "drop" }`
stage.multiline Combine lines stage.multiline { firstline = \^\d{4}` }`
stage.replace Redact/transform stage.replace { expression = \(password=)\S+` replace = "${1}***" }`
stage.structured_metadata Store high-cardinality fields stage.structured_metadata { values = { trace_id = "" } }
stage.tenant Set tenant ID stage.tenant { source = "team" }
stage.cri Parse CRI log format stage.cri {}
stage.sampling Sample a fraction of logs stage.sampling { rate = 0.1 }

Essential CLI Commands

# Query logs via API
curl -G -s "http://loki:3100/loki/api/v1/query_range" \
  --data-urlencode 'query={job="api"} |= "error"' \
  --data-urlencode 'start=1h' \
  | jq

# Check Loki readiness
curl http://loki:3100/ready

# View Loki metrics
curl http://loki:3100/metrics

# Push logs directly
curl -X POST -H "Content-Type: application/json" \
  http://loki:3100/loki/api/v1/push \
  -d '{"streams":[{"stream":{"job":"test"},"values":[["'$(date +%s)000000000'","test log"]]}]}'

# logcli query tool
logcli query '{job="api"}' --limit=100 --since=1h

# Tail logs
logcli query '{job="api"}' --tail

# Check label values
curl http://loki:3100/loki/api/v1/label/job/values

# Alloy: convert an existing Promtail config to Alloy format
alloy convert --source-format=promtail --output=config.alloy promtail.yaml

# Alloy: convert with diagnostic report (warnings about unmappable features)
alloy convert --source-format=promtail \
  --report=report.txt \
  --output=config.alloy promtail.yaml

# Alloy: run directly using a Promtail config (no permanent conversion needed)
alloy run --config.format=promtail promtail.yaml

# Alloy: run with native Alloy config
alloy run config.alloy

# Alloy: check the Alloy UI (targets, component graph, debug info)
# Available at http://alloy-host:12345 by default
curl http://alloy-host:12345/-/ready

Common Issues and Solutions

Issue Cause Solution
"max streams limit exceeded" Too many unique label combinations Reduce label cardinality, remove dynamic labels like pod names or request IDs
Query timeout Query scope too broad or time range too long Add more specific label filters, reduce time range, use split_queries_by_interval
High latency queries Missing index, wide scans Ensure stream selectors are specific, use bloom filters if available
Logs not appearing Alloy not shipping, timestamp issues Check Alloy UI targets at :12345, verify timestamps are within retention window
Out of order entries Clock skew, batching issues Configure max_chunk_age, allow unordered_writes: true
Missing labels Pipeline stage not extracting Check loki.process stage config, verify log format matches stage parser
Duplicate logs Multiple collectors or overlapping targets Check loki.source.file and loki.source.kubernetes target overlap, verify relabelling
High memory usage Too many streams, large chunks Reduce cardinality, tune chunk settings, increase resources
Ingestion rate limited Exceeding limits Increase ingestion_rate_mb, use stage.sampling or stage.drop to reduce volume
Query returns no data Wrong time range or labels Verify labels with /loki/api/v1/labels, check time range
Alloy convert warnings Promtail features with no Alloy equivalent Review --report output; tracing config and some metrics must be manually reconfigured

Debugging Tips

# Alloy UI — component graph, target status, pipeline debug
# Default port: 12345
open http://alloy-host:12345

# Alloy: list active targets (equivalent to Promtail /targets)
curl http://alloy-host:12345/-/ready
# Full component debug via UI at http://alloy-host:12345/graph

# Check label cardinality
curl -s http://loki:3100/loki/api/v1/series \
  --data-urlencode 'match[]={job="api"}' | jq '. | length'

# View ingester ring
curl http://loki:3100/ingester/ring

# Debug log format — add a temporary stage.output to inspect extracted values
# In config.alloy:
#
# loki.process "debug" {
#   forward_to = [loki.write.default.receiver]
#   stage.json {
#     expressions = { level = "level" }
#   }
#   stage.output {
#     source = "level"   // Log line becomes the extracted value for inspection
#   }
# }

# Check chunk utilisation
curl http://loki:3100/metrics | grep loki_ingester_chunk

# Alloy: validate config syntax before deploying
alloy fmt config.alloy          # Format and syntax check

Related Topics

The following topics complement Loki and would enhance your logging and observability capabilities:

  1. Grafana Alloy - The recommended log (and metrics/traces) collection agent for Loki, replacing Promtail
  2. Grafana - Primary visualisation platform for Loki, providing dashboards, exploration, and alerting
  3. Prometheus - Metrics collection system that pairs with Loki for comprehensive observability
  4. OpenTelemetry - Unified collection framework for traces, metrics, and logs that can export to Loki
  5. Kubernetes - Container orchestration commonly monitored with Loki using Alloy DaemonSet deployment
  6. Fluentd/Fluent Bit - Alternative log collectors that can ship to Loki
  7. Jaeger/Tempo - Distributed tracing systems that correlate with Loki logs via trace IDs