Fluentd/Fluent Bit
A unified logging layer for collecting, processing, and routing logs from diverse sources to multiple destinations.
Fluentd/Fluent Bit Cheatsheet
A unified logging layer for collecting, processing, and routing logs from diverse sources to multiple destinations.
Overview
Fluentd and Fluent Bit are open-source data collectors designed for building unified logging infrastructure. Fluentd is feature-rich and extensible, whilst Fluent Bit is lightweight and optimised for containerised environments.
graph LR
subgraph Sources
A[Application Logs]
B[System Logs]
C[Container Logs]
D[Network Devices]
end
subgraph "Fluentd/Fluent Bit"
E[Input Plugins]
F[Parser]
G[Filter Plugins]
H[Buffer]
I[Output Plugins]
end
subgraph Destinations
J[Elasticsearch]
K[S3/Cloud Storage]
L[Kafka]
M[Monitoring Systems]
end
A --> E
B --> E
C --> E
D --> E
E --> F
F --> G
G --> H
H --> I
I --> J
I --> K
I --> L
I --> M
Configuration Syntax
Key Concepts
Fluentd configuration uses a tag-based routing system with three main directive types:
- source: Defines input sources (where logs come from)
- filter: Transforms or enriches log data
- match: Routes processed logs to output destinations
Configuration files use a declarative syntax with directives enclosed in XML-like tags.
graph TD
A[Source] -->|Tagged Events| B[Filter Chain]
B -->|Modified Events| C[Match/Output]
subgraph "Event Structure"
D[Tag: app.logs]
E[Time: 1234567890]
F[Record: JSON data]
end
Common Patterns
Basic Configuration Structure
# Global configuration
<system>
log_level info
workers 4
</system>
# Input source
<source>
@type tail
path /var/log/app/*.log
tag app.logs
<parse>
@type json
</parse>
</source>
# Filter processing
<filter app.**>
@type record_transformer
<record>
hostname "#{Socket.gethostname}"
</record>
</filter>
# Output destination
<match app.**>
@type elasticsearch
host elasticsearch.example.com
port 9200
index_name app-logs
</match>
Fluent Bit Configuration (classic mode / INI format)
Note: This INI-style syntax is Fluent Bit classic mode. As of Fluent Bit v3.2, YAML configuration supports every classic-mode setting plus features classic mode lacks (such as processors), and YAML is now the recommended form. Classic mode is deprecated and scheduled for removal at the end of 2026, so prefer YAML for new configurations. The YAML equivalent of the example below is shown beneath it.
[SERVICE]
Flush 5
Daemon Off
Log_Level info
Parsers_File parsers.conf
[INPUT]
Name tail
Path /var/log/app/*.log
Tag app.logs
Parser json
[FILTER]
Name record_modifier
Match app.*
Record hostname ${HOSTNAME}
[OUTPUT]
Name es
Match app.*
Host elasticsearch.example.com
Port 9200
Index app-logs
The same configuration in YAML (Fluent Bit v3.2+, the current form):
service:
flush: 5
daemon: off
log_level: info
parsers_file: parsers.conf
pipeline:
inputs:
- name: tail
path: /var/log/app/*.log
tag: app.logs
parser: json
filters:
- name: record_modifier
match: app.*
record: hostname ${HOSTNAME}
outputs:
- name: es
match: app.*
host: elasticsearch.example.com
port: 9200
index: app-logs
Examples
Multiple Tags with Label Routing
# Route different log types to separate processing pipelines
<source>
@type tail
path /var/log/nginx/access.log
tag nginx.access
<parse>
@type nginx
</parse>
</source>
<source>
@type tail
path /var/log/nginx/error.log
tag nginx.error
<parse>
@type none
</parse>
</source>
# Use labels for isolated processing
<label @NGINX_PIPELINE>
<filter nginx.**>
@type grep
<regexp>
key code
pattern ^[45]\d{2}$
</regexp>
</filter>
<match nginx.**>
@type elasticsearch
# ... configuration
</match>
</label>
Common Input Plugins
Key Concepts
Input plugins collect logs from various sources and convert them into Fluentd events. Each event consists of:
- Tag: Routing identifier (e.g.,
app.web.access) - Time: Event timestamp
- Record: JSON object containing log data
Common Patterns
tail Plugin (File Monitoring)
<source>
@type tail
path /var/log/app/*.log
pos_file /var/log/fluentd/app.log.pos
tag app.logs
read_from_head true
refresh_interval 5
<parse>
@type json
time_key timestamp
time_format %Y-%m-%dT%H:%M:%S.%NZ
</parse>
</source>
# Multiple file patterns
<source>
@type tail
path /var/log/containers/*.log
exclude_path ["/var/log/containers/fluentd*.log"]
pos_file /var/log/fluentd/containers.pos
tag kubernetes.*
<parse>
@type json
</parse>
</source>
forward Plugin (Agent-to-Aggregator)
# Receiver (Aggregator)
<source>
@type forward
port 24224
bind 0.0.0.0
<security>
self_hostname aggregator.example.com
shared_key secret_key_here
</security>
<transport tls>
cert_path /etc/fluentd/certs/server.crt
private_key_path /etc/fluentd/certs/server.key
</transport>
</source>
# Sender (Agent)
<match **>
@type forward
<server>
host aggregator.example.com
port 24224
</server>
<security>
self_hostname agent01.example.com
shared_key secret_key_here
</security>
</match>
http Plugin (REST API Input)
<source>
@type http
port 8888
bind 0.0.0.0
body_size_limit 32m
keepalive_timeout 10s
<parse>
@type json
</parse>
<transport tls>
cert_path /etc/fluentd/certs/server.crt
private_key_path /etc/fluentd/certs/server.key
</transport>
</source>
# Send logs via HTTP POST
# curl -X POST -d 'json={"message":"test"}' http://localhost:8888/app.logs
Examples
Kubernetes Container Logs
<source>
@type tail
path /var/log/containers/*.log
pos_file /var/log/fluentd/containers.pos
tag kubernetes.*
read_from_head true
<parse>
@type regexp
expression /^(?<time>.+) (?<stream>stdout|stderr) [^ ]* (?<log>.*)$/
time_format %Y-%m-%dT%H:%M:%S.%NZ
</parse>
</source>
Syslog Input
<source>
@type syslog
port 5140
bind 0.0.0.0
tag system
<transport udp>
</transport>
<parse>
@type syslog
with_priority true
</parse>
</source>
Filter Plugins
Key Concepts
Filter plugins modify, enrich, or drop events as they flow through the pipeline. They are processed in the order they appear in the configuration file.
graph LR
A[Input Event] --> B[parser]
B --> C[grep]
C --> D[record_transformer]
D --> E[Output Event]
style B fill:#e1f5fe
style C fill:#fff3e0
style D fill:#f3e5f5
Common Patterns
parser Filter (Re-parse Fields)
# Parse JSON embedded in a string field
<filter app.**>
@type parser
key_name log
reserve_data true
remove_key_name_field true
<parse>
@type json
</parse>
</filter>
# Parse with multiple format support
<filter app.**>
@type parser
key_name message
reserve_data true
<parse>
@type multi_format
<pattern>
format json
</pattern>
<pattern>
format regexp
expression /^(?<level>\w+): (?<message>.*)$/
</pattern>
</parse>
</filter>
grep Filter (Include/Exclude)
# Include only error logs
<filter app.**>
@type grep
<regexp>
key level
pattern /^(error|fatal|critical)$/i
</regexp>
</filter>
# Exclude health check logs
<filter nginx.**>
@type grep
<exclude>
key path
pattern /^\/health(z)?$/
</exclude>
<exclude>
key user_agent
pattern /^kube-probe/
</exclude>
</filter>
# Complex AND/OR conditions
<filter app.**>
@type grep
<and>
<regexp>
key level
pattern error
</regexp>
<regexp>
key service
pattern payment
</regexp>
</and>
</filter>
record_transformer Filter (Modify Records)
# Add static and dynamic fields
<filter app.**>
@type record_transformer
enable_ruby true
<record>
hostname "#{Socket.gethostname}"
environment "#{ENV['ENVIRONMENT'] || 'development'}"
timestamp ${time.strftime('%Y-%m-%dT%H:%M:%S.%LZ')}
tag ${tag}
message_length ${record["message"]&.length || 0}
</record>
</filter>
# Remove sensitive fields
<filter app.**>
@type record_transformer
remove_keys password, credit_card, ssn
</filter>
# Rename fields
<filter app.**>
@type record_transformer
renew_record false
<record>
@timestamp ${record["time"]}
severity ${record["level"]}
</record>
remove_keys time, level
</filter>
Examples
Kubernetes Metadata Enrichment
<filter kubernetes.**>
@type kubernetes_metadata
kubernetes_url https://kubernetes.default.svc
cache_size 1000
watch true
de_dot true
annotation_match ["fluentd.io/.*"]
</filter>
GeoIP Enrichment
<filter nginx.access>
@type geoip
geoip_lookup_keys client_ip
backend_library geoip2_c
<record>
city ${city.names.en["client_ip"]}
country ${country.names.en["client_ip"]}
country_code ${country.iso_code["client_ip"]}
latitude ${location.latitude["client_ip"]}
longitude ${location.longitude["client_ip"]}
</record>
</filter>
Output Plugins
Key Concepts
Output plugins send processed events to external systems. They can be:
- Non-buffered: Send events immediately (e.g., stdout)
- Buffered: Batch events for efficient delivery (e.g., elasticsearch, s3)
- Time-sliced: Organise events by time windows (e.g., file, s3)
Common Patterns
Elasticsearch Output
<match app.**>
@type elasticsearch
host elasticsearch.example.com
port 9200
scheme https
user elastic
password ${ELASTIC_PASSWORD}
# Index configuration
index_name app-logs
# type_name (mapping types) is ignored on Elasticsearch 8+ — mapping
# types were removed in ES 8, so do not set it for modern clusters.
include_timestamp true
# Dynamic index naming
logstash_format true
logstash_prefix app-logs
logstash_dateformat %Y.%m.%d
# Performance tuning
request_timeout 30s
reload_connections false
reconnect_on_error true
reload_on_failure true
# Template management
template_name app-logs
template_file /etc/fluentd/templates/app-logs.json
template_overwrite true
<buffer>
@type file
path /var/log/fluentd/buffer/elasticsearch
flush_mode interval
flush_interval 5s
chunk_limit_size 8MB
queue_limit_length 512
retry_max_interval 30
retry_forever true
</buffer>
</match>
S3 Output
<match logs.**>
@type s3
aws_key_id ${AWS_ACCESS_KEY_ID}
aws_sec_key ${AWS_SECRET_ACCESS_KEY}
s3_bucket my-logs-bucket
s3_region eu-west-1
# Path configuration
path logs/%Y/%m/%d/
s3_object_key_format %{path}%{time_slice}_%{index}.%{file_extension}
# Compression
store_as gzip
# Time slicing
time_slice_format %Y%m%d%H
time_slice_wait 10m
<buffer time>
@type file
path /var/log/fluentd/buffer/s3
timekey 3600
timekey_wait 10m
timekey_use_utc true
chunk_limit_size 256MB
</buffer>
<format>
@type json
</format>
</match>
Kafka Output
<match events.**>
@type kafka2
brokers kafka1:9092,kafka2:9092,kafka3:9092
topic_key topic
default_topic app-events
# Message configuration
use_event_time true
# Authentication
<format>
@type json
</format>
# Security
ssl_ca_cert /etc/fluentd/certs/ca.crt
ssl_client_cert /etc/fluentd/certs/client.crt
ssl_client_cert_key /etc/fluentd/certs/client.key
# SASL authentication
sasl_over_ssl true
username ${KAFKA_USERNAME}
password ${KAFKA_PASSWORD}
scram_mechanism sha256
<buffer topic>
@type file
path /var/log/fluentd/buffer/kafka
flush_interval 3s
chunk_limit_size 1MB
total_limit_size 1GB
</buffer>
</match>
Examples
Multiple Outputs (Copy Plugin)
<match app.**>
@type copy
<store>
@type elasticsearch
host elasticsearch.example.com
# ... elasticsearch config
</store>
<store>
@type s3
s3_bucket backup-logs
# ... s3 config
</store>
<store>
@type stdout
# Debug output
</store>
</match>
Conditional Routing
# Route based on log level
<match app.**>
@type rewrite_tag_filter
<rule>
key level
pattern /^(error|fatal)$/
tag ${tag}.critical
</rule>
<rule>
key level
pattern /^(warn)$/
tag ${tag}.warning
</rule>
<rule>
key level
pattern /.*/
tag ${tag}.info
</rule>
</match>
<match app.**.critical>
@type pagerduty
# Alert configuration
</match>
<match app.**.warning>
@type slack
# Slack notification
</match>
<match app.**.info>
@type elasticsearch
# Standard storage
</match>
Buffering and Retry Strategies
Key Concepts
Buffering ensures reliable log delivery by temporarily storing events before sending them to outputs. Retry mechanisms handle transient failures gracefully.
graph TD
A[Events] --> B[Stage Buffer]
B --> C[Queue Buffer]
C --> D{Flush}
D -->|Success| E[Output]
D -->|Failure| F[Retry Queue]
F -->|Retry| D
F -->|Max Retries| G[Secondary Output]
Common Patterns
Buffer Configuration
<match **>
@type elasticsearch
# ... output config
<buffer tag, time>
@type file
path /var/log/fluentd/buffer/es
# Chunk configuration
chunk_limit_size 8MB
chunk_limit_records 10000
total_limit_size 10GB
# Flush configuration
flush_mode interval
flush_interval 5s
flush_thread_count 4
flush_at_shutdown true
# Retry configuration
retry_type exponential_backoff
retry_wait 1s
retry_max_interval 60s
retry_timeout 72h
retry_forever false
retry_max_times 17
# Overflow behaviour
overflow_action block
# Time-based chunking
timekey 1h
timekey_wait 10m
timekey_use_utc true
</buffer>
</match>
Memory vs File Buffer
# Memory buffer (faster, less durable)
<buffer>
@type memory
chunk_limit_size 8MB
total_limit_size 512MB
flush_interval 1s
retry_max_times 3
</buffer>
# File buffer (slower, more durable)
<buffer>
@type file
path /var/log/fluentd/buffer
chunk_limit_size 8MB
total_limit_size 10GB
flush_interval 5s
retry_forever true
</buffer>
Examples
High-Availability Configuration
<match critical.**>
@type elasticsearch
<buffer>
@type file
path /var/log/fluentd/buffer/critical
# Aggressive flushing for critical logs
flush_mode immediate
flush_thread_count 8
# Extended retry for reliability
retry_forever true
retry_type exponential_backoff
retry_wait 1s
retry_max_interval 300s
# Prevent data loss
overflow_action throw_exception
flush_at_shutdown true
</buffer>
# Secondary output for failures
<secondary>
@type file
path /var/log/fluentd/failed/critical
</secondary>
</match>
Fluent Bit Buffer Configuration
[OUTPUT]
Name es
Match *
Host elasticsearch.example.com
Port 9200
Index logs
# Buffer settings
Buffer_Size 5MB
# Retry settings
Retry_Limit 5
# Storage-based buffering
storage.type filesystem
storage.path /var/log/fluent-bit/buffer/
storage.sync normal
storage.checksum off
storage.max_chunks_up 128
Log Parsing Patterns
Key Concepts
Parsing converts unstructured log text into structured JSON records. Fluentd supports multiple parser types including regexp, json, ltsv, csv, and custom formats.
Common Patterns
Apache/Nginx Access Logs
# Apache combined format
<parse>
@type apache2
</parse>
# Nginx access log
<parse>
@type nginx
</parse>
# Custom nginx format
<parse>
@type regexp
expression /^(?<remote>[^ ]*) (?<host>[^ ]*) (?<user>[^ ]*) \[(?<time>[^\]]*)\] "(?<method>\S+)(?: +(?<path>[^\"]*?)(?: +\S*)?)?" (?<code>[^ ]*) (?<size>[^ ]*)(?: "(?<referer>[^\"]*)" "(?<agent>[^\"]*)")?$/
time_format %d/%b/%Y:%H:%M:%S %z
</parse>
Application Log Patterns
# JSON logs
<parse>
@type json
time_key timestamp
time_format %Y-%m-%dT%H:%M:%S.%NZ
keep_time_key true
</parse>
# Log4j/Logback pattern
<parse>
@type regexp
expression /^(?<time>\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2},\d{3}) (?<level>[A-Z]+) \[(?<thread>[^\]]+)\] (?<class>[^ ]+) - (?<message>.*)$/
time_format %Y-%m-%d %H:%M:%S,%L
</parse>
# Python logging
<parse>
@type regexp
expression /^(?<time>\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2},\d{3}) - (?<name>[^ ]+) - (?<level>[A-Z]+) - (?<message>.*)$/
time_format %Y-%m-%d %H:%M:%S,%L
</parse>
Multi-line Parsing
# Java stack traces
<parse>
@type multiline
format_firstline /^\d{4}-\d{2}-\d{2}/
format1 /^(?<time>\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2},\d{3}) (?<level>[A-Z]+) (?<message>.*)/
</parse>
# Generic multi-line with continuation
<parse>
@type multiline
format_firstline /^[^\s]/
format1 /^(?<message>.*)/
</parse>
Examples
Custom Parser Definitions
# parsers.conf file
<parse>
@type regexp
expression /^\[(?<time>[^\]]+)\] \[(?<level>\w+)\] \[(?<module>[^\]]+)\] (?<message>.*)$/
time_format %Y-%m-%d %H:%M:%S
</parse>
# Kubernetes audit logs
<parse>
@type json
time_key requestReceivedTimestamp
time_format %Y-%m-%dT%H:%M:%S.%NZ
</parse>
Fluent Bit Parsers
# parsers.conf
[PARSER]
Name docker
Format json
Time_Key time
Time_Format %Y-%m-%dT%H:%M:%S.%L
Time_Keep On
[PARSER]
Name syslog-rfc3164
Format regex
Regex /^\<(?<pri>[0-9]+)\>(?<time>[^ ]* {1,2}[^ ]* [^ ]*) (?<host>[^ ]*) (?<ident>[a-zA-Z0-9_\/\.\-]*)(?:\[(?<pid>[0-9]+)\])?(?:[^\:]*\:)? *(?<message>.*)$/
Time_Key time
Time_Format %b %d %H:%M:%S
[PARSER]
Name json
Format json
Time_Key time
Time_Format %d/%b/%Y:%H:%M:%S %z
Quick Reference
| Task | Fluentd Configuration |
|---|---|
| Tail log file | @type tail with path and pos_file |
| Receive forwarded logs | @type forward on port 24224 |
| HTTP input | @type http on specified port |
| Parse JSON | <parse> @type json </parse> |
| Filter by pattern | <filter> @type grep with <regexp> |
| Add fields | <filter> @type record_transformer |
| Output to Elasticsearch | @type elasticsearch with host/index |
| Output to S3 | @type s3 with bucket/region |
| Output to Kafka | @type kafka2 with brokers/topic |
| Buffer to file | <buffer> @type file path /path </buffer> |
| Retry forever | retry_forever true in buffer |
| Copy to multiple outputs | @type copy with multiple <store> |
| Route by tag pattern | <match pattern.**> |
Fluent Bit Quick Reference
| Task | Fluent Bit Configuration |
|---|---|
| Tail log file | [INPUT] Name tail Path /path |
| Parse JSON | Parser json in INPUT section |
| Filter records | [FILTER] Name grep |
| Modify records | [FILTER] Name record_modifier |
| Output to Elasticsearch | [OUTPUT] Name es |
| Output to S3 | [OUTPUT] Name s3 |
| Set log level | Log_Level info in SERVICE |
Common Tag Patterns
| Pattern | Matches |
|---|---|
** |
All tags |
app.** |
Tags starting with app. |
app.{web,api}.** |
Tags starting with app.web. or app.api. |
*.error |
Tags ending with .error |
Common Issues and Solutions
Issue: Logs Not Being Collected
Symptoms: No data appearing in output destination
Solutions:
# Check file permissions
# Fluentd user must have read access to log files
sudo chmod 644 /var/log/app/*.log
sudo chown root:fluentd /var/log/app
# Verify pos_file location is writable
<source>
@type tail
path /var/log/app/*.log
pos_file /var/log/fluentd/app.pos # Must be writable
tag app.logs
</source>
# Enable debug logging
<system>
log_level debug
</system>
# Test configuration
fluentd --dry-run -c /etc/fluentd/fluent.conf
Issue: Buffer Overflow Errors
Symptoms: buffer space has too many data errors
Solutions:
<buffer>
@type file
path /var/log/fluentd/buffer
# Increase buffer limits
total_limit_size 20GB
chunk_limit_size 16MB
# Increase flush frequency
flush_interval 1s
flush_thread_count 8
# Handle overflow gracefully
overflow_action drop_oldest_chunk # or block
</buffer>
Issue: Parsing Failures
Symptoms: Events with unparsed raw logs
Solutions:
# Add error handling for parse failures
<source>
@type tail
path /var/log/app/*.log
tag app.logs
<parse>
@type json
# Emit unparseable lines as raw
@log_level warn
</parse>
</source>
# Use multi-format parser
<parse>
@type multi_format
<pattern>
format json
</pattern>
<pattern>
format none
</pattern>
</parse>
Issue: High Memory Usage
Symptoms: Fluentd consuming excessive memory
Solutions:
# Use file buffer instead of memory
<buffer>
@type file # Not memory
path /var/log/fluentd/buffer
chunk_limit_size 8MB
total_limit_size 1GB
</buffer>
# Limit chunk records
<buffer>
chunk_limit_records 5000
</buffer>
# Flush more frequently
<buffer>
flush_interval 1s
</buffer>
Issue: Elasticsearch Connection Failures
Symptoms: Retry errors, failed to write to Elasticsearch
Solutions:
<match **>
@type elasticsearch
host elasticsearch.example.com
# Increase timeouts
request_timeout 60s
# Enable reconnection
reconnect_on_error true
reload_on_failure true
reload_connections false
# Resilient retry configuration
<buffer>
retry_forever true
retry_type exponential_backoff
retry_wait 1s
retry_max_interval 60s
</buffer>
# Fallback for persistent failures
<secondary>
@type file
path /var/log/fluentd/failed/es
</secondary>
</match>
Issue: Time Zone Mismatches
Symptoms: Incorrect timestamps in output
Solutions:
# Set timezone in system config
<system>
time_format %Y-%m-%dT%H:%M:%S%z
</system>
# Specify timezone in parser
<parse>
@type regexp
expression /(?<time>[^,]+)/
time_format %Y-%m-%d %H:%M:%S
utc true # or local_time true
</parse>
# Use UTC in buffer timekey
<buffer time>
timekey_use_utc true
</buffer>
Issue: Duplicate Logs
Symptoms: Same log entries appearing multiple times
Solutions:
# Ensure unique pos_file per source
<source>
@type tail
path /var/log/app/*.log
pos_file /var/log/fluentd/app-unique.pos # Unique path
tag app.logs
</source>
# Use deduplication filter
<filter app.**>
@type dedot
</filter>
# Check for multiple matching sources
# Ensure tags don't overlap unintentionally
Related Topics
The following topics would complement this Fluentd/Fluent Bit cheatsheet:
-
Elasticsearch - Primary destination for log storage and search; understanding index management, mappings, and query DSL enhances log analysis capabilities
-
Kubernetes Logging - Container orchestration logging patterns, including DaemonSet deployment of Fluent Bit, namespace-based routing, and metadata enrichment
-
Prometheus and Grafana - Monitoring Fluentd/Fluent Bit metrics, creating dashboards for pipeline health, and alerting on log processing issues
-
Kafka - Message queue integration for high-throughput log streaming, understanding topics, partitions, and consumer groups for log pipelines
-
Regular Expressions - Essential for creating custom parsers, grep filters, and rewrite rules; mastering regex improves log parsing accuracy
-
OpenTelemetry - Modern observability framework that complements Fluentd for unified telemetry collection including logs, metrics, and traces