Loki
Log aggregation system designed for efficiency and cost-effectiveness, using labels for indexing rather than full-text indexing.
Loki
Log aggregation system designed for efficiency and cost-effectiveness, using labels for indexing rather than full-text indexing.
Overview
Loki is a horizontally scalable, highly available, multi-tenant log aggregation system inspired by Prometheus. Unlike traditional log systems that index the full content of logs, Loki only indexes metadata (labels), making it significantly cheaper to operate and easier to scale. It integrates seamlessly with Grafana and uses LogQL, a query language similar to PromQL, for querying logs.
graph TB
subgraph "Loki Architecture"
A[Applications] -->|logs| B[Alloy/Agents]
B -->|push| C[Distributor]
C --> D[Ingester]
D --> E[(Object Storage)]
D --> F[(Index Store)]
G[Query Frontend] --> H[Querier]
H --> D
H --> E
H --> F
I[Grafana] --> G
end
subgraph "Log Sources"
J[Docker]
K[Kubernetes]
L[Syslog]
M[Files]
end
J --> A
K --> A
L --> A
M --> A
LogQL Query Language
LogQL is Loki's query language, combining log stream selection with filtering and aggregation capabilities.
Key Concepts
flowchart LR
A[Stream Selector] --> B[Line Filters]
B --> C[Parser]
C --> D[Label Filters]
D --> E[Metric Queries]
subgraph "Query Pipeline"
A
B
C
D
E
end
Query Types:
- Log queries: Return log lines matching criteria
- Metric queries: Calculate values from log content over time
Stream Selectors use labels to filter log streams, similar to Prometheus label selectors.
Stream Selectors
# Exact match
{job="nginx"}
# Regex match
{job=~"nginx|apache"}
# Not equal
{job!="debug"}
# Negative regex match
{job!~"test.*"}
# Multiple labels
{job="api", environment="production", level="error"}
# Combine matchers
{namespace=~"prod-.*", container!="sidecar"}
Line Filters
# Contains string (case-sensitive)
{job="nginx"} |= "error"
# Does not contain
{job="nginx"} != "debug"
# Regex match
{job="nginx"} |~ "status=[45][0-9]{2}"
# Negative regex match
{job="nginx"} !~ "health_check"
# Chain multiple filters (AND logic)
{job="api"} |= "error" != "timeout" |~ "user_id=[0-9]+"
# Case-insensitive match
{job="nginx"} |~ "(?i)error"
Parsers
# JSON parser - extracts all JSON fields as labels
{job="api"} | json
# Extract specific JSON fields
{job="api"} | json level, message, user_id
# Logfmt parser
{job="api"} | logfmt
# Regex parser with named groups
{job="nginx"} | regexp `(?P<ip>\d+\.\d+\.\d+\.\d+) - - \[(?P<timestamp>[^\]]+)\] "(?P<method>\w+) (?P<path>[^"]+)"`
# Pattern parser (simplified regex)
{job="nginx"} | pattern `<ip> - - [<timestamp>] "<method> <path> <_>" <status> <size>`
# Unpack parser for logs packed with stage.pack (Alloy) or Promtail pack stage
{job="api"} | unpack
# Line format - restructure log line
{job="api"} | json | line_format "{{.level}} - {{.message}}"
Label Filters
# Filter on extracted labels
{job="api"} | json | level="error"
# Numeric comparisons
{job="api"} | json | status >= 400
# Duration comparisons
{job="api"} | json | response_time > 500ms
# Byte comparisons
{job="api"} | json | size > 1kb
# Combine filters
{job="api"} | json | level="error" and status >= 500
# OR logic
{job="api"} | json | level="error" or level="warn"
# Regex on extracted labels
{job="api"} | json | message =~ ".*timeout.*"
Metric Queries
# Count logs per second over time
rate({job="nginx"}[5m])
# Count specific errors
rate({job="nginx"} |= "error" [1m])
# Sum by label
sum by (level) (rate({job="api"} | json [5m]))
# Count unique values
count_over_time({job="api"}[1h])
# Bytes rate
bytes_rate({job="nginx"}[5m])
# Bytes over time
bytes_over_time({job="nginx"}[1h])
# Quantile from extracted values
quantile_over_time(0.95, {job="api"} | json | unwrap response_time [5m])
# Average of extracted numeric field
avg_over_time({job="api"} | json | unwrap duration [5m]) by (endpoint)
# Absent over time (detect missing logs)
absent_over_time({job="critical-service"}[5m])
Advanced Aggregations
# Top 5 paths by request count
topk(5, sum by (path) (rate({job="nginx"} | pattern `<_> "<_> <path> <_>"` [5m])))
# Error rate percentage
sum(rate({job="api"} | json | level="error" [5m]))
/ sum(rate({job="api"} [5m])) * 100
# Aggregate and sort by multiple dimensions
sum by (service, level) (count_over_time({namespace="production"} | json [1h]))
# Compare with offset
sum(rate({job="api"} |= "error" [1h]))
- sum(rate({job="api"} |= "error" [1h] offset 24h))
# Group left/right for joining
sum by (instance) (rate({job="api"} |= "error" [5m]))
/ on (instance) group_left
sum by (instance) (rate({job="api"} [5m]))
Examples
# Find all 5xx errors in nginx logs
{job="nginx"} |~ "HTTP/[0-9.]+ [5][0-9]{2}"
# Parse structured API logs and filter
{job="api", environment="production"}
| json
| level="error"
| line_format "{{.timestamp}} [{{.level}}] {{.message}}"
# Calculate p99 response time from JSON logs
quantile_over_time(0.99,
{job="api"}
| json
| unwrap response_ms
| __error__=""
[5m]
) by (endpoint)
# Count errors by service and error type
sum by (service, error_type) (
count_over_time(
{namespace="production"}
| json
| level="error"
[1h]
)
)
# Find slow requests
{job="api"}
| json
| response_time > 1s
| line_format "{{.method}} {{.path}} took {{.response_time}}"
Label Strategies
Effective label design is crucial for Loki performance and query efficiency.
Key Concepts
graph TB
subgraph "Good Labels"
A[Static values]
B[Low cardinality]
C[Useful for filtering]
end
subgraph "Bad Labels"
D[Dynamic values]
E[High cardinality]
F[Unique identifiers]
end
A --> G[Good Performance]
B --> G
C --> G
D --> H[Poor Performance]
E --> H
F --> H
Cardinality Guidelines:
- Keep total unique label combinations under 10,000 per tenant
- Each unique label set creates a new stream
- High cardinality causes increased memory usage and query times
Common Label Patterns
# Recommended labels
labels:
job: nginx # Application name
environment: production # Environment
namespace: web-services # Kubernetes namespace
cluster: eu-west-1 # Cluster/region
level: error # Log level (if consistent)
# Avoid these as labels
labels:
request_id: "abc-123" # High cardinality - unique per request
user_id: "12345" # High cardinality - unique per user
timestamp: "2024-01-01T..." # Already stored in log entry
ip_address: "10.0.0.1" # High cardinality - many unique IPs
trace_id: "xyz-789" # High cardinality - unique per trace
Label Design Patterns
# Pattern 1: Environment-based hierarchy
labels:
environment: production
region: eu-west-1
cluster: main
namespace: api
service: user-service
pod: user-service-abc123 # Only if truly needed
# Pattern 2: Team ownership
labels:
team: platform
service: auth
component: api
# Pattern 3: Log categorisation
labels:
job: application
level: info # Only if log level is static per stream
format: json # Helps choose parser
# Pattern 4: Kubernetes standard labels
labels:
namespace: default
deployment: my-app
container: main
pod: my-app-xyz # Consider if this adds value
Dynamic Labelling Strategies
// Alloy: extract level as label using loki.process (use cautiously — high cardinality risk)
loki.process "api_labels" {
forward_to = [loki.write.default.receiver]
stage.json {
expressions = { level = "level" }
}
stage.labels {
values = { level = "" }
}
}
loki.source.file "api" {
targets = [{
__path__ = "/var/log/api/*.log",
job = "api",
}]
forward_to = [loki.process.api_labels.receiver]
}
// Drop high-cardinality labels after discovery
loki.process "nginx_drop" {
forward_to = [loki.write.default.receiver]
stage.label_drop {
values = ["pod_template_hash", "controller_revision_hash"]
}
}
Stream Optimisation
// Consolidate similar streams
// Instead of labels: {app="api", instance="pod-1"}, {app="api", instance="pod-2"} …
// Consider: {app="api"}
// Use structured metadata or line content for per-instance identification.
// Alloy: keep essential labels only via discovery.relabel
discovery.relabel "k8s_essential" {
targets = discovery.kubernetes.pods.targets
rule {
action = "labelkeep"
regex = "(namespace|container|pod)"
}
rule {
source_labels = ["__meta_kubernetes_namespace"]
target_label = "namespace"
}
}
Grafana Alloy — Log Collection Agent
Grafana Alloy is the recommended log collection agent for Loki. It supersedes Promtail, which entered Long-Term Support on 13 February 2025 and reached End-of-Life on 2 March 2026. All new deployments should use Alloy.
Migration:
alloy convert --source-format=promtail --output=config.alloy promtail.yamlconverts an existing Promtail config to Alloy format automatically. You can also run Alloy directly against a Promtail config file withalloy run --config.format=promtail promtail.yamlwhile you migrate. See the official migration guide.Legacy Promtail: if you still operate Promtail in a legacy environment, the pipeline-stage concepts (JSON, regex, timestamp, labels, multiline, drop, match, replace, pack, tenant) map directly to Alloy
stage.*blocks insideloki.process. The syntax differs but the semantics are identical.
Key Concepts
flowchart TB
A[Service Discovery] --> B[Target Files / K8s Pods]
B --> C[loki.source.file / loki.source.kubernetes]
C --> D[loki.process — Pipeline Stages]
D --> E[loki.write — Push to Loki]
subgraph "Pipeline Stages (loki.process)"
F[stage.json / stage.regex] --> G[stage.timestamp]
G --> H[stage.labels / stage.static_labels]
H --> I[stage.drop / stage.match]
I --> J[stage.output / stage.replace]
end
D --> F
Core Alloy components for log collection:
local.file_match: Discovers files on disk using glob patternsloki.source.file: Tails discovered files and forwards log entriesloki.source.kubernetes: Tails Kubernetes pod logs via the API (no node privileges needed)discovery.kubernetes: Discovers Kubernetes resources for targetingdiscovery.relabel: Applies relabelling rules to discovered targetsloki.process: Applies processing stages (parse, transform, filter)loki.write: Ships logs to a Loki endpoint
Basic File Collection
// config.alloy — minimal file collection
// Discover log files
local.file_match "system" {
path_targets = [
{ __path__ = "/var/log/*.log", job = "varlogs" },
]
}
// Tail matched files and forward to Loki
loki.source.file "system" {
targets = local.file_match.system.targets
forward_to = [loki.write.default.receiver]
}
// Multiple path patterns
local.file_match "applications" {
path_targets = [
{ __path__ = "/var/log/apps/**/*.log", job = "apps" },
]
}
loki.source.file "applications" {
targets = local.file_match.applications.targets
forward_to = [loki.write.default.receiver]
}
// Loki endpoint
loki.write "default" {
endpoint {
url = "http://loki:3100/loki/api/v1/push"
tenant_id = "default"
}
external_labels = {
cluster = "production-eu",
}
}
Processing Stages (loki.process)
// JSON API logs: parse, timestamp, label, output, redact
local.file_match "api" {
path_targets = [{
__path__ = "/var/log/api/*.log",
job = "api",
environment = "production",
}]
}
loki.source.file "api" {
targets = local.file_match.api.targets
forward_to = [loki.process.api.receiver]
}
loki.process "api" {
forward_to = [loki.write.default.receiver]
// Parse JSON log line
stage.json {
expressions = {
level = "level",
message = "message",
ts = "timestamp",
duration = "response_time",
}
}
// Set timestamp from extracted field
// Supported shorthand: RFC3339, RFC3339Nano, Unix, UnixMs, UnixUs, UnixNs
stage.timestamp {
source = "ts"
format = "RFC3339Nano"
}
// Promote extracted fields to labels (low-cardinality only)
stage.labels {
values = { level = "" }
}
// Add static labels
stage.static_labels {
values = { environment = "production" }
}
// Rewrite the log line to just the message field
stage.output {
source = "message"
}
}
// Nginx access log: regex parsing
local.file_match "nginx" {
path_targets = [{
__path__ = "/var/log/nginx/access.log",
job = "nginx",
}]
}
loki.source.file "nginx" {
targets = local.file_match.nginx.targets
forward_to = [loki.process.nginx.receiver]
}
loki.process "nginx" {
forward_to = [loki.write.default.receiver]
// Use backtick raw strings for regex — no double-escaping needed
stage.regex {
expression = `^(?P<ip>\S+) - (?P<user>\S+) \[(?P<ts>[^\]]+)\] "(?P<method>\S+) (?P<path>\S+) (?P<proto>[^"]+)" (?P<status>\d+) (?P<size>\d+)`
}
stage.timestamp {
source = "ts"
format = "02/Jan/2006:15:04:05 -0700"
}
// Derive status group (2xx, 4xx, 5xx) with template
stage.template {
source = "status_group"
template = `{{ .status | substr 0 1 }}xx`
}
stage.labels {
values = { status_group = "" }
}
}
Advanced Processing Stages
loki.process "advanced" {
forward_to = [loki.write.default.receiver]
// Multiline logs — collapse stack traces into a single entry
stage.multiline {
firstline = `^\d{4}-\d{2}-\d{2}`
max_wait_time = "3s"
max_lines = 128
}
// Drop health-check noise
stage.match {
selector = `{job="nginx"} |~ ".*health_check.*"`
action = "drop"
drop_counter_reason = "health_check"
}
// Derive metrics from log content
stage.metrics {
metric.counter {
name = "http_requests_total"
description = "Total HTTP requests"
prefix = "alloy_custom_"
source = "status"
action = "inc"
}
metric.histogram {
name = "request_duration_seconds"
description = "Request duration"
prefix = "alloy_custom_"
source = "duration"
buckets = [0.01, 0.05, 0.1, 0.5, 1, 5]
}
}
// Redact sensitive data
stage.replace {
expression = `(password=)[^&\s]+`
replace = "${1}***REDACTED***"
}
// Pack high-cardinality fields into the log line body instead of labels
stage.pack {
labels = ["level", "service"]
ingest_timestamp = true
}
// Set tenant from an extracted label
stage.tenant {
source = "team"
}
}
Kubernetes Log Collection
// Pod logs via the Kubernetes API — no DaemonSet node privileges required.
// Use loki.source.kubernetes for API-based tailing,
// or loki.source.file (DaemonSet) for file-based tailing from /var/log/pods.
// --- Option A: API-based (recommended for most clusters) ---
discovery.kubernetes "pods" {
role = "pod"
// When running as a DaemonSet, restrict to pods on this node
selectors {
role = "pod"
field = "spec.nodeName=" + coalesce(sys.env("HOSTNAME"), constants.hostname)
}
}
discovery.relabel "pod_logs" {
targets = discovery.kubernetes.pods.targets
// Only scrape pods with the opt-in annotation
rule {
source_labels = ["__meta_kubernetes_pod_annotation_alloy_io_scrape"]
action = "keep"
regex = "true"
}
rule {
source_labels = ["__meta_kubernetes_namespace"]
target_label = "namespace"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_pod_container_name"]
target_label = "container"
}
rule {
source_labels = ["__meta_kubernetes_pod_label_app_kubernetes_io_name"]
target_label = "app"
}
rule {
source_labels = ["__meta_kubernetes_namespace", "__meta_kubernetes_pod_container_name"]
separator = "/"
replacement = "$1"
target_label = "job"
}
// Drop verbose Kubernetes meta-labels from the final label set
rule {
action = "labeldrop"
regex = "__meta_kubernetes_pod_label_(.+)"
}
}
loki.source.kubernetes "pod_logs" {
targets = discovery.relabel.pod_logs.output
forward_to = [loki.process.pod_logs.receiver]
}
loki.process "pod_logs" {
forward_to = [loki.write.default.receiver]
// CRI format is standard for containerd / CRI-O runtimes
stage.cri {}
// Add cluster-level static label
stage.static_labels {
values = { cluster = "production-eu" }
}
}
// --- Option B: file-based DaemonSet (requires host /var/log/pods mount) ---
discovery.relabel "pod_file_logs" {
targets = discovery.kubernetes.pods.targets
rule {
source_labels = ["__meta_kubernetes_pod_uid", "__meta_kubernetes_pod_container_name"]
separator = "/"
replacement = "/var/log/pods/*$1/*.log"
target_label = "__path__"
}
rule {
source_labels = ["__meta_kubernetes_namespace"]
target_label = "namespace"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_pod_container_name"]
target_label = "container"
}
}
loki.source.file "pod_file_logs" {
targets = discovery.relabel.pod_file_logs.output
forward_to = [loki.process.pod_logs.receiver]
}
Complete Production Example
// config.alloy — production application logs
// --- Loki endpoint ---
loki.write "default" {
endpoint {
url = "http://loki-gateway:3100/loki/api/v1/push"
// For Grafana Cloud / authenticated endpoints:
// basic_auth {
// username = "<USERNAME>"
// password = "<API_KEY>"
// }
}
external_labels = {
cluster = "production-eu",
}
}
// --- Application logs with multiline + JSON pipeline ---
local.file_match "application" {
path_targets = [{
__path__ = "/var/log/app/*.log",
job = "application",
environment = "production",
}]
}
loki.source.file "application" {
targets = local.file_match.application.targets
forward_to = [loki.process.application.receiver]
}
loki.process "application" {
forward_to = [loki.write.default.receiver]
// Collapse multiline entries (ISO 8601 timestamp starts a new entry)
stage.multiline {
firstline = `^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}`
max_wait_time = "3s"
}
stage.json {
expressions = {
level = "level",
service = "service",
trace_id = "trace_id",
ts = "@timestamp",
}
}
stage.timestamp {
source = "ts"
format = "RFC3339Nano"
}
stage.labels {
values = {
level = "",
service = "",
}
}
// trace_id is high-cardinality — store as structured metadata, not a label
stage.structured_metadata {
values = { trace_id = "" }
}
// Drop debug logs in production
stage.match {
selector = `{level="debug"}`
action = "drop"
drop_counter_reason = "debug_logs"
}
}
// --- Kubernetes cluster events ---
loki.source.kubernetes_events "cluster_events" {
job_name = "integrations/kubernetes/eventhandler"
log_format = "logfmt"
forward_to = [loki.process.cluster_events.receiver]
}
loki.process "cluster_events" {
forward_to = [loki.write.default.receiver]
stage.static_labels {
values = { cluster = "production-eu" }
}
}
Integration with Grafana
Grafana provides the primary visualisation interface for Loki, enabling log exploration, dashboards, and alerting.
Key Concepts
sequenceDiagram
participant U as User
participant G as Grafana
participant L as Loki
participant S as Storage
U->>G: Query logs
G->>L: LogQL query
L->>S: Fetch chunks
S-->>L: Log data
L->>L: Process & filter
L-->>G: Results
G-->>U: Visualisation
Integration Features:
- Explore view for ad-hoc queries
- Dashboard panels for visualisation
- Alerting on log patterns
- Log context and correlations
Data Source Configuration
# Grafana datasource provisioning
apiVersion: 1
datasources:
- name: Loki
type: loki
access: proxy
url: http://loki:3100
isDefault: false
jsonData:
maxLines: 1000
derivedFields:
# Link to traces
- name: TraceID
matcherRegex: '"trace_id":"(\w+)"'
url: '$${__value.raw}'
datasourceUid: tempo
urlDisplayLabel: View Trace
# Link to external system
- name: RequestID
matcherRegex: 'request_id=(\w+)'
url: 'https://logs.example.com/request/$${__value.raw}'
urlDisplayLabel: View in External System
secureJsonData:
# For multi-tenant setups
httpHeaderValue1: tenant-id
Dashboard Panels
{
"panels": [
{
"title": "Error Logs",
"type": "logs",
"datasource": "Loki",
"targets": [
{
"expr": "{job=\"api\"} |= \"error\" | json",
"refId": "A"
}
],
"options": {
"showTime": true,
"showLabels": true,
"wrapLogMessage": true,
"prettifyLogMessage": true,
"enableLogDetails": true,
"dedupStrategy": "none",
"sortOrder": "Descending"
}
},
{
"title": "Log Volume",
"type": "timeseries",
"datasource": "Loki",
"targets": [
{
"expr": "sum by (level) (rate({job=\"api\"} | json [5m]))",
"legendFormat": "{{level}}",
"refId": "A"
}
]
},
{
"title": "Error Rate",
"type": "stat",
"datasource": "Loki",
"targets": [
{
"expr": "sum(rate({job=\"api\"} |= \"error\" [5m])) / sum(rate({job=\"api\"} [5m])) * 100",
"refId": "A"
}
],
"options": {
"colorMode": "value",
"graphMode": "none"
},
"fieldConfig": {
"defaults": {
"unit": "percent",
"thresholds": {
"mode": "absolute",
"steps": [
{"color": "green", "value": null},
{"color": "yellow", "value": 1},
{"color": "red", "value": 5}
]
}
}
}
}
]
}
Alerting Rules
# Loki ruler configuration
groups:
- name: application-alerts
interval: 1m
rules:
# Alert on error rate
- alert: HighErrorRate
expr: |
sum(rate({job="api"} |= "error" [5m]))
/ sum(rate({job="api"} [5m])) > 0.05
for: 5m
labels:
severity: warning
annotations:
summary: "High error rate in API logs"
description: "Error rate is {{ $value | humanizePercentage }}"
# Alert on specific error pattern
- alert: DatabaseConnectionError
expr: |
count_over_time({job="api"} |= "database connection failed" [5m]) > 5
for: 2m
labels:
severity: critical
annotations:
summary: "Database connection errors detected"
description: "{{ $value }} connection errors in last 5 minutes"
# Alert on missing logs
- alert: NoLogsReceived
expr: |
absent_over_time({job="critical-service"}[5m])
for: 5m
labels:
severity: critical
annotations:
summary: "No logs from critical service"
description: "No logs received from critical-service for 5 minutes"
Correlating with Metrics and Traces
# Grafana data source correlations
datasources:
- name: Loki
type: loki
jsonData:
derivedFields:
# Trace correlation
- name: traceID
matcherRegex: '"traceID":"(\w+)"'
url: '$${__value.raw}'
datasourceUid: tempo
# Link to metrics
- name: service
matcherRegex: '"service":"(\w+)"'
url: '/explore?left={"datasource":"prometheus","queries":[{"expr":"up{service=\"$${__value.raw}\"}"}]}'
urlDisplayLabel: View Metrics
Explore View Tips
# Use query builder for complex queries
# Start with stream selector
{job="api", namespace="production"}
# Add filters progressively
{job="api"} |= "error"
# Parse and filter
{job="api"} | json | level="error" | status >= 500
# Add time range in URL
/explore?left={"range":{"from":"now-1h","to":"now"},"queries":[...]}
# Split view for comparison
# Use right panel for different query or time range
Log Aggregation Patterns
Design patterns for collecting, routing, and managing logs at scale.
Key Concepts
graph TB
subgraph "Collection Patterns"
A[Direct Push]
B[Sidecar]
C[DaemonSet]
D[Aggregator]
end
subgraph "Routing"
E[By Label]
F[By Content]
G[By Tenant]
end
subgraph "Storage"
H[Hot]
I[Warm]
J[Cold]
end
A --> E
B --> E
C --> E
D --> E
E --> H
F --> H
G --> H
H --> I
I --> J
Collection Architectures
# Pattern 1: DaemonSet (Kubernetes) — Alloy per node, file-based log access
# Install via the official Grafana Alloy Helm chart:
# helm upgrade --install alloy grafana/alloy -f alloy-values.yaml
# The chart mounts /var/log and /var/lib/docker/containers automatically.
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: alloy
spec:
template:
spec:
containers:
- name: alloy
image: grafana/alloy:latest
args:
- run
- /etc/alloy/config.alloy
volumeMounts:
- name: config
mountPath: /etc/alloy
- name: varlog
mountPath: /var/log
readOnly: true
- name: containers
mountPath: /var/lib/docker/containers
readOnly: true
volumes:
- name: config
configMap:
name: alloy-config
- name: varlog
hostPath:
path: /var/log
- name: containers
hostPath:
path: /var/lib/docker/containers
---
# Pattern 2: Sidecar (per-pod) — Alloy sidecar for apps with custom log paths
apiVersion: v1
kind: Pod
metadata:
name: app-with-alloy-sidecar
spec:
containers:
- name: app
image: myapp
volumeMounts:
- name: logs
mountPath: /var/log/app
- name: alloy
image: grafana/alloy:latest
args:
- run
- /etc/alloy/config.alloy
volumeMounts:
- name: logs
mountPath: /var/log/app
readOnly: true
- name: config
mountPath: /etc/alloy
volumes:
- name: logs
emptyDir: {}
- name: config
configMap:
name: alloy-sidecar-config
---
# Pattern 3: Aggregator — Alloy fan-out to multiple Loki clusters
# In config.alloy, define multiple loki.write endpoints:
#
# loki.write "primary" {
# endpoint { url = "http://loki-primary:3100/loki/api/v1/push" }
# }
# loki.write "secondary" {
# endpoint { url = "http://loki-secondary:3100/loki/api/v1/push" }
# external_labels = { replica = "secondary" }
# }
#
# Then forward_to = [loki.write.primary.receiver, loki.write.secondary.receiver]
Multi-Tenant Patterns
// Alloy: static tenant ID on the loki.write endpoint
loki.write "team_a" {
endpoint {
url = "http://loki:3100/loki/api/v1/push"
tenant_id = "team-a"
}
external_labels = { tenant = "team-a" }
}
// Dynamic tenant ID from an extracted label
loki.process "multi_tenant" {
forward_to = [loki.write.team_a.receiver]
stage.json {
expressions = { team = "team" }
}
// Sets the X-Scope-OrgID header from the extracted 'team' value
stage.tenant {
source = "team"
}
}
# Loki server: multi-tenant configuration
auth_enabled: true
limits_config:
ingestion_rate_mb: 10
ingestion_burst_size_mb: 20
max_streams_per_user: 10000
max_global_streams_per_user: 50000
Log Routing and Filtering
// Alloy: route, filter, and sample logs based on content
loki.process "router" {
// Fan-out: send to both primary and secondary clusters simultaneously
forward_to = [
loki.write.primary.receiver,
loki.write.secondary.receiver,
]
stage.json {
expressions = {
level = "level",
service = "service",
}
}
// Route: promote level and service as labels for stream selection
stage.labels {
values = {
level = "",
service = "",
}
}
// Drop debug logs in production
stage.match {
selector = `{level="debug"}`
action = "drop"
drop_counter_reason = "debug_dropped"
}
// Sample verbose info logs — keep 10%
stage.match {
selector = `{level="info"}`
stage.sampling {
rate = 0.1
drop_counter_reason = "info_sampled"
}
}
}
loki.write "primary" {
endpoint {
url = "http://loki-primary:3100/loki/api/v1/push"
}
}
loki.write "secondary" {
endpoint {
url = "http://loki-secondary:3100/loki/api/v1/push"
}
external_labels = { replica = "secondary" }
}
Retention and Lifecycle
# Loki retention configuration
# Loki 3.0 defaults to the TSDB index and schema v13; boltdb-shipper
# and schema v11 are legacy and not recommended for new deployments.
schema_config:
configs:
- from: 2024-01-01
store: tsdb
object_store: s3
schema: v13
index:
prefix: loki_index_
period: 24h
compactor:
working_directory: /data/loki/compactor
# `shared_store` was removed in Loki 3.0. The compactor now uses the
# object store from storage_config; set delete_request_store (required
# when retention is enabled) to your object store.
delete_request_store: s3
compaction_interval: 10m
retention_enabled: true
retention_delete_delay: 2h
retention_delete_worker_count: 150
limits_config:
retention_period: 720h # 30 days default
# Per-stream retention
overrides:
tenant-a:
retention_period: 2160h # 90 days for tenant-a
tenant-b:
retention_period: 168h # 7 days for tenant-b
Performance Optimisation
Strategies for improving Loki query performance, ingestion throughput, and resource efficiency.
Key Concepts
graph LR
subgraph "Performance Factors"
A[Query Scope] --> D[Performance]
B[Label Design] --> D
C[Chunk Config] --> D
end
subgraph "Optimisation Areas"
E[Ingestion]
F[Storage]
G[Querying]
end
D --> E
D --> F
D --> G
Performance Principles:
- Minimise label cardinality
- Use appropriate chunk sizing
- Cache effectively
- Parallelise queries
Query Optimisation
# Bad: No stream selector
|= "error"
# Good: Specific stream selector
{job="api", namespace="production"} |= "error"
# Bad: Wide time range
{job="api"} # defaults to last hour or more
# Good: Narrow time range
{job="api"} # with explicit short time range in query
# Bad: Inefficient filter order
{job="api"} |~ "user_id=\d+" |= "error"
# Good: Most selective filter first
{job="api"} |= "error" |~ "user_id=\d+"
# Bad: Unnecessary parsing
{job="api"} | json | line_format "{{.message}}"
# Good: Only parse if needed for filtering
{job="api"} |= "error" | json | level="error"
# Use bloom filters (if enabled)
{job="api"} |= "error" # Exact match uses bloom filter
Ingestion Optimisation
# Loki ingester configuration
ingester:
chunk_idle_period: 30m # Flush inactive chunks
chunk_block_size: 262144 # 256KB uncompressed
chunk_target_size: 1572864 # 1.5MB compressed target
chunk_retain_period: 30s # Keep chunks in memory after flush
max_transfer_retries: 0 # Disable transfers in StatefulSet
wal:
enabled: true
dir: /data/loki/wal
flush_on_shutdown: true
replay_memory_ceiling: 4GB
# Distributor configuration
distributor:
ring:
kvstore:
store: memberlist
# Limits for ingestion
limits_config:
ingestion_rate_mb: 20
ingestion_burst_size_mb: 30
per_stream_rate_limit: 5MB
per_stream_rate_limit_burst: 15MB
max_line_size: 256kb
max_entries_limit_per_query: 5000
Storage Optimisation
# Chunk storage configuration
# TSDB is the recommended index store from Loki 2.8+ (default in 3.0).
# The `shared_store` field was removed in Loki 3.0; the shipper now
# derives the object store from the storage_config backend below.
storage_config:
tsdb_shipper:
active_index_directory: /data/loki/tsdb-index
cache_location: /data/loki/tsdb-cache
cache_ttl: 24h
aws:
s3: s3://region/bucket-name
s3forcepathstyle: false
# Filesystem for development
filesystem:
directory: /data/loki/chunks
# Caching
chunk_store_config:
max_look_back_period: 0s
chunk_cache_config:
embedded_cache:
enabled: true
max_size_mb: 500
ttl: 1h
query_range:
results_cache:
cache:
embedded_cache:
enabled: true
max_size_mb: 100
# Index caching
index_queries_cache_config:
embedded_cache:
enabled: true
max_size_mb: 100
Query Frontend Optimisation
# Query frontend configuration
frontend:
max_outstanding_per_tenant: 2048
compress_responses: true
log_queries_longer_than: 5s
query_scheduler:
max_outstanding_requests_per_tenant: 2048
querier:
max_concurrent: 10
query_ingesters_within: 3h # Only query ingesters for recent data
query_range:
align_queries_with_step: true
max_retries: 5
cache_results: true
parallelise_shardable_queries: true
# Split queries by time interval
limits_config:
split_queries_by_interval: 30m
max_query_parallelism: 32
max_query_length: 721h # 30 days
max_query_series: 500
Resource Tuning
# Kubernetes resource configuration
resources:
# Ingester - memory intensive
ingester:
requests:
cpu: "1"
memory: 4Gi
limits:
cpu: "2"
memory: 8Gi
# Querier - CPU intensive
querier:
requests:
cpu: "2"
memory: 2Gi
limits:
cpu: "4"
memory: 4Gi
# Query Frontend - moderate resources
query-frontend:
requests:
cpu: "500m"
memory: 1Gi
limits:
cpu: "1"
memory: 2Gi
# JVM tuning for components using BoltDB
# Environment variables
GOGC: "80" # Tune garbage collection
GOMEMLIMIT: "6GiB" # Soft memory limit
Monitoring Loki Performance
# Ingestion rate
sum(rate(loki_distributor_bytes_received_total[5m]))
# Query latency
histogram_quantile(0.99, sum(rate(loki_request_duration_seconds_bucket{route=~"loki_api_v1_query.*"}[5m])) by (le))
# Chunk flush rate
sum(rate(loki_ingester_chunks_flushed_total[5m]))
# Cache hit rate
sum(rate(loki_cache_hits_total[5m])) / sum(rate(loki_cache_fetched_keys_total[5m]))
# Stream count per tenant
loki_ingester_streams_created_total - loki_ingester_streams_removed_total
# Query queue length
loki_query_scheduler_queue_length
Quick Reference
LogQL Syntax
| Element | Syntax | Example |
|---|---|---|
| Stream selector | {label="value"} |
{job="nginx"} |
| Exact match | = |
{env="prod"} |
| Regex match | =~ |
{job=~"api|web"} |
| Not equal | != |
{env!="dev"} |
| Contains | |= |
{job="api"} |= "error" |
| Not contains | != |
{job="api"} != "debug" |
| Regex line | |~ |
{job="api"} |~ "status=[45].." |
| JSON parser | | json |
{job="api"} | json |
| Label filter | | field="value" |
{job="api"} | json | level="error" |
| Rate | rate({...}[5m]) |
rate({job="api"}[5m]) |
| Count | count_over_time({...}[1h]) |
count_over_time({job="api"} |= "error"[1h]) |
Common LogQL Patterns
| Pattern | Query |
|---|---|
| Error count | count_over_time({job="api"} |= "error" [1h]) |
| Error rate | rate({job="api"} |= "error" [5m]) |
| Top errors by type | topk(10, sum by (error) (count_over_time({job="api"} | json [1h]))) |
| Log volume by level | sum by (level) (rate({job="api"} | json [5m])) |
| P99 latency | quantile_over_time(0.99, {job="api"} | json | unwrap duration [5m]) |
| Missing logs | absent_over_time({job="critical"}[5m]) |
Alloy Processing Stages (loki.process)
| Stage block | Purpose | Example |
|---|---|---|
stage.json |
Parse JSON | stage.json { expressions = { level = "level" } } |
stage.logfmt |
Parse logfmt | stage.logfmt { mapping = { level = "" } } |
stage.regex |
Parse with regex | stage.regex { expression = \(?P<ip>\S+)` }` |
stage.timestamp |
Set timestamp | stage.timestamp { source = "ts" format = "RFC3339" } |
stage.labels |
Add dynamic labels | stage.labels { values = { level = "" } } |
stage.static_labels |
Add static labels | stage.static_labels { values = { env = "prod" } } |
stage.output |
Rewrite log line | stage.output { source = "message" } |
stage.drop |
Drop lines | stage.drop { expression = ".*health.*" } |
stage.match |
Conditional stages | stage.match { selector = \{job="api"}` action = "drop" }` |
stage.multiline |
Combine lines | stage.multiline { firstline = \^\d{4}` }` |
stage.replace |
Redact/transform | stage.replace { expression = \(password=)\S+` replace = "${1}***" }` |
stage.structured_metadata |
Store high-cardinality fields | stage.structured_metadata { values = { trace_id = "" } } |
stage.tenant |
Set tenant ID | stage.tenant { source = "team" } |
stage.cri |
Parse CRI log format | stage.cri {} |
stage.sampling |
Sample a fraction of logs | stage.sampling { rate = 0.1 } |
Essential CLI Commands
# Query logs via API
curl -G -s "http://loki:3100/loki/api/v1/query_range" \
--data-urlencode 'query={job="api"} |= "error"' \
--data-urlencode 'start=1h' \
| jq
# Check Loki readiness
curl http://loki:3100/ready
# View Loki metrics
curl http://loki:3100/metrics
# Push logs directly
curl -X POST -H "Content-Type: application/json" \
http://loki:3100/loki/api/v1/push \
-d '{"streams":[{"stream":{"job":"test"},"values":[["'$(date +%s)000000000'","test log"]]}]}'
# logcli query tool
logcli query '{job="api"}' --limit=100 --since=1h
# Tail logs
logcli query '{job="api"}' --tail
# Check label values
curl http://loki:3100/loki/api/v1/label/job/values
# Alloy: convert an existing Promtail config to Alloy format
alloy convert --source-format=promtail --output=config.alloy promtail.yaml
# Alloy: convert with diagnostic report (warnings about unmappable features)
alloy convert --source-format=promtail \
--report=report.txt \
--output=config.alloy promtail.yaml
# Alloy: run directly using a Promtail config (no permanent conversion needed)
alloy run --config.format=promtail promtail.yaml
# Alloy: run with native Alloy config
alloy run config.alloy
# Alloy: check the Alloy UI (targets, component graph, debug info)
# Available at http://alloy-host:12345 by default
curl http://alloy-host:12345/-/ready
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| "max streams limit exceeded" | Too many unique label combinations | Reduce label cardinality, remove dynamic labels like pod names or request IDs |
| Query timeout | Query scope too broad or time range too long | Add more specific label filters, reduce time range, use split_queries_by_interval |
| High latency queries | Missing index, wide scans | Ensure stream selectors are specific, use bloom filters if available |
| Logs not appearing | Alloy not shipping, timestamp issues | Check Alloy UI targets at :12345, verify timestamps are within retention window |
| Out of order entries | Clock skew, batching issues | Configure max_chunk_age, allow unordered_writes: true |
| Missing labels | Pipeline stage not extracting | Check loki.process stage config, verify log format matches stage parser |
| Duplicate logs | Multiple collectors or overlapping targets | Check loki.source.file and loki.source.kubernetes target overlap, verify relabelling |
| High memory usage | Too many streams, large chunks | Reduce cardinality, tune chunk settings, increase resources |
| Ingestion rate limited | Exceeding limits | Increase ingestion_rate_mb, use stage.sampling or stage.drop to reduce volume |
| Query returns no data | Wrong time range or labels | Verify labels with /loki/api/v1/labels, check time range |
| Alloy convert warnings | Promtail features with no Alloy equivalent | Review --report output; tracing config and some metrics must be manually reconfigured |
Debugging Tips
# Alloy UI — component graph, target status, pipeline debug
# Default port: 12345
open http://alloy-host:12345
# Alloy: list active targets (equivalent to Promtail /targets)
curl http://alloy-host:12345/-/ready
# Full component debug via UI at http://alloy-host:12345/graph
# Check label cardinality
curl -s http://loki:3100/loki/api/v1/series \
--data-urlencode 'match[]={job="api"}' | jq '. | length'
# View ingester ring
curl http://loki:3100/ingester/ring
# Debug log format — add a temporary stage.output to inspect extracted values
# In config.alloy:
#
# loki.process "debug" {
# forward_to = [loki.write.default.receiver]
# stage.json {
# expressions = { level = "level" }
# }
# stage.output {
# source = "level" // Log line becomes the extracted value for inspection
# }
# }
# Check chunk utilisation
curl http://loki:3100/metrics | grep loki_ingester_chunk
# Alloy: validate config syntax before deploying
alloy fmt config.alloy # Format and syntax check
Related Topics
The following topics complement Loki and would enhance your logging and observability capabilities:
- Grafana Alloy - The recommended log (and metrics/traces) collection agent for Loki, replacing Promtail
- Grafana - Primary visualisation platform for Loki, providing dashboards, exploration, and alerting
- Prometheus - Metrics collection system that pairs with Loki for comprehensive observability
- OpenTelemetry - Unified collection framework for traces, metrics, and logs that can export to Loki
- Kubernetes - Container orchestration commonly monitored with Loki using Alloy DaemonSet deployment
- Fluentd/Fluent Bit - Alternative log collectors that can ship to Loki
- Jaeger/Tempo - Distributed tracing systems that correlate with Loki logs via trace IDs