etcd Operations
Essential commands and patterns for managing etcd clusters in production environments, with focus on Kubernetes deployments.
etcd Operations
Essential commands and patterns for managing etcd clusters in production environments, with focus on Kubernetes deployments.
Overview
etcd is a distributed, strongly consistent key-value store that provides reliable storage for critical distributed system data. It uses the Raft consensus algorithm to maintain consistency across cluster members and is the primary datastore for Kubernetes cluster state.
graph TB
subgraph etcd Cluster
subgraph Member 1 - Leader
L[Leader Node<br/>Accepts R/W]
LS[(Local Storage)]
end
subgraph Member 2 - Follower
F1[Follower Node<br/>Accepts R]
F1S[(Local Storage)]
end
subgraph Member 3 - Follower
F2[Follower Node<br/>Accepts R]
F2S[(Local Storage)]
end
end
subgraph Clients
K8S[Kubernetes<br/>API Server]
CLI[etcdctl CLI]
APP[Other Apps]
end
K8S -->|Write| L
CLI -->|Read| F1
APP -->|Read| F2
L <-->|Raft<br/>Replication| F1
L <-->|Raft<br/>Replication| F2
F1 <-->|Raft<br/>Heartbeat| F2
L --> LS
F1 --> F1S
F2 --> F2S
Cluster Topology and Quorum
Understanding cluster member roles and quorum requirements for high availability.
Key Concepts
flowchart TB
subgraph Quorum Decision
N[Number of Members: N]
Q[Quorum Required: N/2 + 1]
F[Maximum Failures: floor N-1/2]
end
subgraph Examples
E3[3 Members<br/>Quorum: 2<br/>Tolerate: 1 failure]
E5[5 Members<br/>Quorum: 3<br/>Tolerate: 2 failures]
E7[7 Members<br/>Quorum: 4<br/>Tolerate: 3 failures]
end
N --> Q
Q --> F
F --> E3
F --> E5
F --> E7
- Cluster Size: Always use odd numbers (3, 5, 7); even numbers don't improve fault tolerance
- Quorum: Majority of members must agree for writes; calculated as
(N/2) + 1 - Leader Election: Uses Raft consensus; members vote to elect a leader
- Follower: Forwards write requests to leader; serves read requests
- Learner: Non-voting member useful for scaling reads without affecting quorum
Cluster Member Management
# List cluster members
etcdctl member list
etcdctl member list --write-out=table
# Check endpoint status
etcdctl endpoint status --cluster --write-out=table
# Check endpoint health
etcdctl endpoint health --cluster
# Add a new member (returns peer URLs to configure on new node)
etcdctl member add node4 \
--peer-urls=https://10.0.1.14:2380
# Remove a member (use ID from member list)
etcdctl member remove 8e9e05c52164694d
# Update member peer URLs
etcdctl member update 8e9e05c52164694d \
--peer-urls=https://10.0.1.11:2380
# Add a learner (non-voting member)
etcdctl member add node4 \
--learner \
--peer-urls=https://10.0.1.14:2380
# Promote learner to voting member
etcdctl member promote 8e9e05c52164694d
Cluster Health Monitoring
# Check overall cluster health
etcdctl endpoint health \
--endpoints=https://10.0.1.11:2379,https://10.0.1.12:2379,https://10.0.1.13:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# Get detailed endpoint status
etcdctl endpoint status --cluster -w table
# Monitor cluster alarms
etcdctl alarm list
# Disarm all alarms (after fixing underlying issues)
etcdctl alarm disarm
Examples
# Bootstrap a new 3-node cluster
# Node 1
etcd --name=node1 \
--initial-advertise-peer-urls=https://10.0.1.11:2380 \
--listen-peer-urls=https://10.0.1.11:2380 \
--listen-client-urls=https://10.0.1.11:2379,https://127.0.0.1:2379 \
--advertise-client-urls=https://10.0.1.11:2379 \
--initial-cluster-token=etcd-cluster-1 \
--initial-cluster=node1=https://10.0.1.11:2380,node2=https://10.0.1.12:2380,node3=https://10.0.1.13:2380 \
--initial-cluster-state=new
# Add a learner to scale reads without affecting quorum
etcdctl member add node4 --learner --peer-urls=https://10.0.1.14:2380
# Start node4 with --initial-cluster-state=existing
# Wait for sync to complete
etcdctl member promote <member-id>
Compaction and Defragmentation
Managing historical revisions and reclaiming storage space.
Key Concepts
flowchart LR
subgraph etcd Storage
direction TB
KV[Key-Value Store]
REV[Revision History]
TOMB[Tombstones]
end
subgraph Operations
COMP[Compaction<br/>Remove old revisions]
DEFRAG[Defragmentation<br/>Reclaim disk space]
end
REV --> COMP
COMP --> DEFRAG
DEFRAG --> KV
subgraph Before
B1[DB Size: 10GB<br/>Used: 3GB<br/>Fragmented: 7GB]
end
subgraph After
A1[DB Size: 3GB<br/>Used: 3GB<br/>Fragmented: 0GB]
end
DEFRAG --> B1
B1 --> A1
- Revisions: etcd keeps history of all changes; each modification creates a new revision
- Compaction: Removes old revisions to save space; compaction is not reversible
- Auto-Compaction: Automatic periodic compaction based on time or revision count
- Defragmentation: Reclaims storage space after compaction; requires per-member execution
- MVCC: Multi-version concurrency control enables watch and historical queries
Compaction Operations
# Get current revision number
etcdctl endpoint status --write-out=table
# Manual compaction to specific revision
etcdctl compact 1000000
# Check compacted revision
etcdctl endpoint status -w table | grep -i compact
# Enable auto-compaction (retention mode: periodic - 1 hour)
etcd --auto-compaction-mode=periodic \
--auto-compaction-retention=1h
# Enable auto-compaction (revision mode - keep last 1000 revisions)
etcd --auto-compaction-mode=revision \
--auto-compaction-retention=1000
Defragmentation Operations
# Defragment all cluster members (recommended)
etcdctl defrag --cluster
# Defragment specific endpoint
etcdctl defrag --endpoints=https://10.0.1.11:2379
# Check DB size before and after defragmentation
etcdctl endpoint status --cluster -w table
# Automated defragmentation script for all members
for endpoint in https://10.0.1.11:2379 https://10.0.1.12:2379 https://10.0.1.13:2379; do
echo "Defragmenting $endpoint..."
etcdctl defrag --endpoints=$endpoint \
--command-timeout=30s || echo "Failed to defrag $endpoint"
sleep 5
done
Examples
# Kubernetes cluster maintenance workflow
# 1. Check current database size
etcdctl endpoint status --cluster -w table \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# 2. Get current revision
CURRENT_REV=$(etcdctl endpoint status --write-out=json | \
jq -r '.[] | .Status.header.revision')
# 3. Compact to current revision minus 1000
etcdctl compact $((CURRENT_REV - 1000))
# 4. Defragment each member sequentially
etcdctl defrag --cluster
# 5. Verify space reclaimed
etcdctl endpoint status --cluster -w table
# Configure auto-compaction in Kubernetes static pod manifest
# /etc/kubernetes/manifests/etcd.yaml
# Add to command args:
# - --auto-compaction-mode=periodic
# - --auto-compaction-retention=5m
Backup and Restore
Critical procedures for disaster recovery and cluster migration.
Key Concepts
flowchart TB
subgraph Backup Strategy
direction LR
SNAP[Snapshot<br/>Point-in-time backup]
CRON[Scheduled Backups<br/>Automated snapshots]
OFF[Offsite Storage<br/>S3/GCS/NFS]
end
subgraph Restore Process
direction TB
STOP[Stop all etcd members]
REST[Restore from snapshot<br/>on each node]
START[Start cluster with<br/>existing state]
end
SNAP --> CRON
CRON --> OFF
OFF --> STOP
STOP --> REST
REST --> START
subgraph Verification
CHECK[Verify data integrity]
VAL[Validate cluster health]
end
START --> CHECK
CHECK --> VAL
- Snapshot: Complete point-in-time backup of etcd data; includes all keys and metadata
- Consistency: Snapshots are consistent; capture cluster state at specific revision
- Restore: Requires stopping all members; creates new cluster from snapshot
- Data Directory: Never manually copy data directory while etcd is running
- etcdutl vs etcdctl (3.5+): the live
snapshot savestays inetcdctl(it talks to a running server), but the offlinesnapshot statusandsnapshot restoremoved to theetcdutlbinary.etcdctlstill accepts them but prints a deprecation notice.
Backup Procedures
# Create snapshot backup
etcdctl snapshot save /backup/etcd-snapshot-$(date +%Y%m%d-%H%M%S).db
# With Kubernetes etcd (using host path)
etcdctl snapshot save /var/lib/etcd-backup/snapshot.db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# Verify snapshot integrity (offline op -> etcdutl in 3.5+)
etcdutl snapshot status /backup/etcd-snapshot.db -w table
# Get snapshot metadata
etcdutl snapshot status /backup/etcd-snapshot.db -w json | jq
# Automated backup script
#!/bin/bash
BACKUP_DIR="/var/lib/etcd-backup"
RETENTION_DAYS=7
TIMESTAMP=$(date +%Y%m%d-%H%M%S)
# Create backup
etcdctl snapshot save ${BACKUP_DIR}/snapshot-${TIMESTAMP}.db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# Verify backup
etcdutl snapshot status ${BACKUP_DIR}/snapshot-${TIMESTAMP}.db
# Remove old backups
find ${BACKUP_DIR} -name "snapshot-*.db" -mtime +${RETENTION_DAYS} -delete
# Upload to S3 (optional)
aws s3 cp ${BACKUP_DIR}/snapshot-${TIMESTAMP}.db \
s3://my-bucket/etcd-backups/
Restore Procedures
# Restore from snapshot to new data directory
etcdutl snapshot restore /backup/etcd-snapshot.db \
--data-dir=/var/lib/etcd-restored \
--name=node1 \
--initial-cluster=node1=https://10.0.1.11:2380,node2=https://10.0.1.12:2380,node3=https://10.0.1.13:2380 \
--initial-cluster-token=etcd-cluster-restored \
--initial-advertise-peer-urls=https://10.0.1.11:2380
# Restore on all cluster members (each node runs its own restore)
# Node 1
etcdutl snapshot restore /backup/snapshot.db \
--name=node1 \
--initial-cluster=node1=https://10.0.1.11:2380,node2=https://10.0.1.12:2380,node3=https://10.0.1.13:2380 \
--initial-cluster-token=etcd-cluster-1 \
--initial-advertise-peer-urls=https://10.0.1.11:2380 \
--data-dir=/var/lib/etcd
# Node 2
etcdutl snapshot restore /backup/snapshot.db \
--name=node2 \
--initial-cluster=node1=https://10.0.1.11:2380,node2=https://10.0.1.12:2380,node3=https://10.0.1.13:2380 \
--initial-cluster-token=etcd-cluster-1 \
--initial-advertise-peer-urls=https://10.0.1.12:2380 \
--data-dir=/var/lib/etcd
# Node 3
etcdutl snapshot restore /backup/snapshot.db \
--name=node3 \
--initial-cluster=node1=https://10.0.1.11:2380,node2=https://10.0.1.12:2380,node3=https://10.0.1.13:2380 \
--initial-cluster-token=etcd-cluster-1 \
--initial-advertise-peer-urls=https://10.0.1.13:2380 \
--data-dir=/var/lib/etcd
Examples
# Kubernetes etcd disaster recovery workflow
# 1. Create emergency backup (if cluster is still accessible)
kubectl exec -n kube-system etcd-master-1 -- sh -c \
"ETCDCTL_API=3 etcdctl snapshot save /var/lib/etcd/backup.db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key"
# 2. Copy backup from pod to host
kubectl cp -n kube-system etcd-master-1:/var/lib/etcd/backup.db \
/root/etcd-backup.db
# 3. Stop kubelet on all control plane nodes
systemctl stop kubelet
# 4. Move static pod manifests (stops etcd)
mv /etc/kubernetes/manifests/etcd.yaml /tmp/
# 5. Backup current data directory
mv /var/lib/etcd /var/lib/etcd.old
# 6. Restore snapshot on each control plane node
etcdutl snapshot restore /root/etcd-backup.db \
--data-dir=/var/lib/etcd \
--name=$(hostname) \
--initial-cluster=master-1=https://10.0.1.11:2380,master-2=https://10.0.1.12:2380,master-3=https://10.0.1.13:2380 \
--initial-cluster-token=etcd-cluster \
--initial-advertise-peer-urls=https://$(hostname -i):2380
# 7. Set correct permissions
chown -R etcd:etcd /var/lib/etcd
# 8. Restore static pod manifest
mv /tmp/etcd.yaml /etc/kubernetes/manifests/
# 9. Start kubelet
systemctl start kubelet
# 10. Verify cluster health
etcdctl endpoint health --cluster
kubectl get nodes
Authentication and RBAC
Securing etcd with authentication and role-based access control.
Key Concepts
flowchart TB
subgraph Authentication
ROOT[Root User<br/>Full access]
USERS[Regular Users<br/>Assigned roles]
end
subgraph RBAC
ROLES[Roles<br/>Permission sets]
PERMS[Permissions<br/>Read/Write/RW]
end
subgraph Authorization
KEYSPACE[Key Ranges<br/>/kubernetes/*<br/>/app/*]
end
ROOT --> ROLES
USERS --> ROLES
ROLES --> PERMS
PERMS --> KEYSPACE
- Authentication: Verify client identity using username/password or client certificates
- Roles: Named collections of permissions; assigned to users
- Permissions: Read, Write, or ReadWrite on specific key ranges
- Root User: Special user with full cluster access; required for RBAC setup
- Client Certificates: TLS-based authentication; stronger than password-based
Enabling Authentication
# Add root user (required before enabling auth)
etcdctl user add root
# Enter password when prompted
# The root user must also hold the built-in "root" role, or auth enable
# warns ("root user does not have root role") and may be rejected
etcdctl role add root
etcdctl user grant-role root root
# Enable authentication
etcdctl auth enable
# After enabling auth, all commands require authentication
export ETCDCTL_USER=root:password
# Or specify with each command
etcdctl --user=root:password member list
# Disable authentication (emergency only)
etcdctl --user=root:password auth disable
User Management
# Create a new user
etcdctl user add myuser
etcdctl --user=root:password user add myuser
# List all users
etcdctl user list
# Get user details
etcdctl user get myuser
# Grant role to user
etcdctl user grant-role myuser myrole
# Revoke role from user
etcdctl user revoke-role myuser myrole
# Change user password
etcdctl user passwd myuser
# Delete user
etcdctl user delete myuser
Role Management
# Create a role
etcdctl role add readonly
# Grant read permission to a role
etcdctl role grant-permission readonly read /app/
# Grant read permission to specific key range
etcdctl role grant-permission readonly read /app/config /app/configz
# Grant write permission
etcdctl role grant-permission readwrite write /app/
# Grant readwrite permission
etcdctl role grant-permission readwrite readwrite /app/
# List all roles
etcdctl role list
# Get role details and permissions
etcdctl role get readonly
# Revoke permission from role
etcdctl role revoke-permission readonly /app/
# Delete role
etcdctl role delete readonly
Examples
# Kubernetes RBAC setup for external clients
# 1. Enable authentication (as root)
etcdctl user add root
etcdctl auth enable
# 2. Create role for Kubernetes API server
etcdctl --user=root:password role add k8s-apiserver
# 3. Grant full access to Kubernetes keyspace
etcdctl --user=root:password role grant-permission k8s-apiserver \
readwrite /registry/
# 4. Create user for API server
etcdctl --user=root:password user add apiserver
# 5. Grant role to user
etcdctl --user=root:password user grant-role apiserver k8s-apiserver
# Create read-only monitoring user
etcdctl --user=root:password role add monitoring
etcdctl --user=root:password role grant-permission monitoring read ""
etcdctl --user=root:password user add prometheus
etcdctl --user=root:password user grant-role prometheus monitoring
# Application-specific access
etcdctl --user=root:password role add app-config-reader
etcdctl --user=root:password role grant-permission app-config-reader \
read /app/production/config/
etcdctl --user=root:password user add config-service
etcdctl --user=root:password user grant-role config-service app-config-reader
# Verify permissions
etcdctl --user=config-service:password get /app/production/config/database
Performance Tuning and Client Settings
Optimising etcd performance for production workloads.
Key Concepts
flowchart TB
subgraph Server Performance
direction LR
DISK[Disk I/O<br/>SSDs recommended<br/>Latency critical]
NET[Network Latency<br/>RTT affects writes<br/>Dedicated NICs]
MEM[Memory<br/>8GB+ recommended<br/>Watch cache]
end
subgraph Client Optimisation
direction LR
POOL[Connection Pooling<br/>Reuse connections]
TIMEOUT[Timeouts<br/>Appropriate values]
RETRY[Retry Logic<br/>Exponential backoff]
end
subgraph Limits
direction LR
QUOTA[Storage Quota<br/>Default: 2GB]
SIZE[Request Size<br/>Max: 1.5MB]
WATCH[Watch Streams<br/>Limit active watches]
end
DISK --> QUOTA
NET --> TIMEOUT
MEM --> WATCH
- Disk Latency: Critical for performance; use SSDs with low latency (<10ms)
- Network RTT: Affects consensus; deploy members in same datacenter
- Storage Quota: Default 2GB; increase for large clusters (Kubernetes typically needs 4-8GB)
- Request Size: Single request limited to 1.5MB; use pagination for large results
- Watch Channels: Memory-intensive; limit concurrent watches
Server Configuration
# Set storage quota (default: 2GB)
etcd --quota-backend-bytes=8589934592 # 8GB
# Snapshot settings for performance
etcd --snapshot-count=10000 # Trigger snapshot every 10k transactions
# Heartbeat and election timeouts
etcd --heartbeat-interval=100 \
--election-timeout=1000
# Optimise for write-heavy workloads
etcd --max-txn-ops=10240 \
--max-request-bytes=10485760 # 10MB
# Enable metrics for monitoring
etcd --listen-metrics-urls=http://0.0.0.0:2381
# Disk priority (ionice) for etcd process
ionice -c2 -n0 -p $(pgrep etcd)
# CPU priority (nice) for etcd process
renice -n -10 -p $(pgrep etcd)
Client Configuration
# Set environment variables for optimal client behaviour
export ETCDCTL_DIAL_TIMEOUT=10s
export ETCDCTL_COMMAND_TIMEOUT=30s
export ETCDCTL_KEEPALIVE_TIME=10s
export ETCDCTL_KEEPALIVE_TIMEOUT=3s
# Connection settings in code (example with Go client)
# Config{
# Endpoints: []string{"https://10.0.1.11:2379"},
# DialTimeout: 10 * time.Second,
# DialKeepAliveTime: 10 * time.Second,
# DialKeepAliveTimeout: 3 * time.Second,
# MaxCallSendMsgSize: 2 * 1024 * 1024, // 2MB
# MaxCallRecvMsgSize: 4 * 1024 * 1024, // 4MB
# }
# Use pagination for large result sets
etcdctl get /registry/pods --prefix --limit=100 --keys-only
# Serialisable reads for lower latency (no linearisable guarantee)
etcdctl get /mykey --consistency=s
# Watch with filters to reduce overhead
etcdctl watch /myprefix --prefix --rev=1000000
Monitoring Metrics
# Key metrics to monitor
curl http://localhost:2381/metrics | grep -E \
'etcd_server_proposals_failed_total|etcd_server_leader_changes_seen_total|etcd_disk_wal_fsync_duration_seconds|etcd_disk_backend_commit_duration_seconds|etcd_mvcc_db_total_size_in_bytes'
# Check backend database size
etcdctl endpoint status --cluster -w table
# Monitor slow operations
# Watch for wal_fsync > 10ms or backend_commit > 25ms
# Prometheus queries for alerting
# etcd_disk_wal_fsync_duration_seconds_bucket{le="0.01"} < 0.99
# etcd_disk_backend_commit_duration_seconds_bucket{le="0.025"} < 0.99
# etcd_mvcc_db_total_size_in_bytes > 6000000000 # 6GB threshold
Examples
# Kubernetes etcd performance tuning
# 1. Increase storage quota in static pod manifest
# /etc/kubernetes/manifests/etcd.yaml
# spec:
# containers:
# - command:
# - etcd
# - --quota-backend-bytes=8589934592 # 8GB
# - --snapshot-count=10000
# 2. Monitor database size
watch -n 60 'etcdctl endpoint status --cluster -w table'
# 3. Set up alerts for space quota
etcdctl alarm list
# If NOSPACE alarm is active:
# - Take snapshot backup
# - Compact old revisions
# - Defragment cluster
# - Increase quota if needed
# 4. Optimise Kubernetes API server flags
# /etc/kubernetes/manifests/kube-apiserver.yaml
# - --etcd-servers=https://127.0.0.1:2379
# - --etcd-compaction-interval=5m # Auto-compact every 5 minutes
# 5. Test etcd performance
# Write performance
benchmark --endpoints=https://127.0.0.1:2379 \
--conns=100 --clients=1000 \
put --key-size=8 --sequential-keys \
--total=100000 --val-size=256
# Read performance
benchmark --endpoints=https://127.0.0.1:2379 \
--conns=100 --clients=1000 \
range /test --total=100000
Quick Reference
Essential Commands
| Command | Description |
|---|---|
etcdctl member list |
List all cluster members |
etcdctl endpoint health --cluster |
Check health of all endpoints |
etcdctl endpoint status --cluster -w table |
Get detailed status in table format |
etcdctl snapshot save backup.db |
Create backup snapshot |
etcdutl snapshot restore backup.db |
Restore from snapshot |
etcdctl compact <revision> |
Compact history to revision |
etcdctl defrag --cluster |
Defragment all members |
etcdctl alarm list |
Check for active alarms |
etcdctl user add <username> |
Create new user |
etcdctl role add <rolename> |
Create new role |
etcdctl auth enable |
Enable authentication |
etcdctl get <key> --prefix |
Get all keys with prefix |
etcdctl put <key> <value> |
Set key-value pair |
etcdctl del <key> --prefix |
Delete keys with prefix |
etcdctl watch <key> --prefix |
Watch for changes |
Environment Variables
| Variable | Purpose |
|---|---|
ETCDCTL_API=3 |
Select v3 API (default since etcd 3.4; only needed for older clients) |
ETCDCTL_ENDPOINTS |
Comma-separated endpoint URLs |
ETCDCTL_CACERT |
CA certificate file path |
ETCDCTL_CERT |
Client certificate file path |
ETCDCTL_KEY |
Client private key file path |
ETCDCTL_USER |
Username:password for auth |
ETCDCTL_DIAL_TIMEOUT |
Connection timeout (default: 2s) |
ETCDCTL_COMMAND_TIMEOUT |
Command timeout (default: 5s) |
Kubernetes etcd Locations
| Distribution | etcd Manifest Path | Data Directory |
|---|---|---|
| kubeadm | /etc/kubernetes/manifests/etcd.yaml |
/var/lib/etcd |
| kops | Managed via instance groups | /mnt/master-vol-*/var/etcd |
| EKS | Managed by AWS (no direct access) | N/A |
| GKE | Managed by Google (no direct access) | N/A |
| AKS | Managed by Azure (no direct access) | N/A |
Recommended Cluster Sizes
| Cluster Size | Quorum | Max Failures | Use Case |
|---|---|---|---|
| 1 | 1 | 0 | Development only |
| 3 | 2 | 1 | Small production |
| 5 | 3 | 2 | Large production |
| 7 | 4 | 3 | Mission-critical |
Storage Quota Recommendations
| Environment | Quota | Notes |
|---|---|---|
| Development | 2GB | Default |
| Small K8s (<50 nodes) | 4GB | Typical workload |
| Medium K8s (50-200 nodes) | 8GB | Heavy workload |
| Large K8s (200+ nodes) | 16GB | Very heavy workload |
Common Issues and Solutions
Space Quota Exceeded
# Symptoms
etcdctl alarm list
# Output: memberID:xxxxx alarm:NOSPACE
# Solution
# 1. Check database size
etcdctl endpoint status --cluster -w table
# 2. Take backup first
etcdctl snapshot save /backup/emergency-backup.db
# 3. Compact old revisions
CURRENT_REV=$(etcdctl endpoint status --write-out=json | \
jq -r '.[] | .Status.header.revision')
etcdctl compact $((CURRENT_REV - 1000))
# 4. Defragment to reclaim space
etcdctl defrag --cluster
# 5. Disarm alarms
etcdctl alarm disarm
# 6. If still insufficient, increase quota
# Edit /etc/kubernetes/manifests/etcd.yaml
# Add: --quota-backend-bytes=8589934592
Slow Performance / High Latency
# Diagnosis
# Check disk latency
iostat -x 5
# Check etcd metrics
curl http://localhost:2381/metrics | grep fsync_duration
# Solutions
# 1. Verify SSD usage (not HDD)
lsblk -d -o name,rota
# rota=0 means SSD, rota=1 means HDD
# 2. Disable disk barriers if using dedicated disk
# Add to etcd systemd service:
# Environment="ETCD_UNSUPPORTED_ARCH=x86_64"
# 3. Check for CPU throttling
# Increase CPU priority
renice -n -10 -p $(pgrep etcd)
# 4. Reduce heartbeat interval for faster leader election
# Edit etcd config:
# --heartbeat-interval=100
# --election-timeout=1000
# 5. Enable metrics and monitor
# --listen-metrics-urls=http://0.0.0.0:2381
Member Unhealthy / Lost Quorum
# Check cluster health
etcdctl endpoint health --cluster
# If majority of members are down (lost quorum)
# 1. Verify network connectivity between members
ping <member-ip>
nc -zv <member-ip> 2380
# 2. Check etcd logs for errors
journalctl -u etcd -f
# or for Kubernetes:
kubectl logs -n kube-system etcd-master-1
# 3. If member is permanently lost, remove and re-add
# From healthy member:
etcdctl member list
etcdctl member remove <member-id>
etcdctl member add <member-name> --peer-urls=https://<new-ip>:2380
# 4. If quorum is lost, disaster recovery needed
# Restore from backup on all members (see Backup and Restore section)
Data Inconsistency
# Check for hash mismatches
etcdctl endpoint hashkv --cluster
# If hashes don't match
# 1. Identify the outlier member
# 2. Remove the member
etcdctl member remove <member-id>
# 3. Delete its data directory
rm -rf /var/lib/etcd
# 4. Re-add as new member
etcdctl member add <name> --peer-urls=https://<ip>:2380
# 5. Start etcd with --initial-cluster-state=existing
TLS Certificate Issues
# Verify certificate validity
openssl x509 -in /etc/kubernetes/pki/etcd/server.crt -text -noout
# Check certificate expiry
openssl x509 -in /etc/kubernetes/pki/etcd/server.crt -noout -enddate
# Verify certificate chain
openssl verify -CAfile /etc/kubernetes/pki/etcd/ca.crt \
/etc/kubernetes/pki/etcd/server.crt
# Test TLS connection
openssl s_client -connect 127.0.0.1:2379 \
-cert /etc/kubernetes/pki/etcd/server.crt \
-key /etc/kubernetes/pki/etcd/server.key \
-CAfile /etc/kubernetes/pki/etcd/ca.crt
# For Kubernetes, certificates can be renewed with:
kubeadm certs check-expiration
kubeadm certs renew etcd-server
kubeadm certs renew etcd-peer
High Watch Load
# Check active watchers
etcdctl watch --prefix / --rev=0 &
# Monitor memory usage
watch -n 5 'ps aux | grep etcd'
# Solutions
# 1. Reduce number of watchers in applications
# 2. Use filtering in watches
# 3. Increase server resources (especially memory)
# 4. Monitor with:
curl http://localhost:2381/metrics | grep etcd_debugging_mvcc_watcher_total
# 5. Cap concurrent gRPC streams per client (HTTP/2; bounds watch/lease fan-out)
# --max-concurrent-streams=1000
Database Size Growing Rapidly
# Identify large keys
etcdctl get --prefix / --keys-only | while read key; do
size=$(etcdctl get "$key" --print-value-only | wc -c)
echo "$size $key"
done | sort -rn | head -20
# Enable auto-compaction
# In etcd config or Kubernetes API server:
# --auto-compaction-mode=periodic
# --auto-compaction-retention=5m
# For Kubernetes, also configure in kube-apiserver:
# --etcd-compaction-interval=5m
# Schedule regular defragmentation (cron)
0 2 * * * /usr/local/bin/etcdctl defrag --cluster >> /var/log/etcd-defrag.log 2>&1
Leader Election Failures
# Check leader election metrics
curl http://localhost:2381/metrics | grep leader_changes
# Frequent leader changes indicate:
# 1. Network issues between members
# 2. Resource starvation (CPU/disk)
# 3. Incorrect timeouts
# Solutions
# 1. Check network latency between members
ping -c 100 <member-ip> | tail -1
# 2. Increase election timeout
# --election-timeout=5000 # 5 seconds
# 3. Ensure stable network
# Use dedicated network interfaces for etcd
# Avoid cross-region cluster deployment
# 4. Monitor with:
etcdctl endpoint status --cluster -w table
# Look for Leader column changing frequently