Kubernetes Troubleshooting Playbook
Systematic approaches for diagnosing and resolving common Kubernetes cluster issues.
Kubernetes Troubleshooting Playbook
Systematic approaches for diagnosing and resolving common Kubernetes cluster issues.
Overview
Kubernetes troubleshooting requires understanding the interaction between control plane components, worker nodes, networking layers, and storage systems. This playbook provides structured diagnostic workflows for identifying and resolving issues across pods, nodes, networking, storage, and resource management. Each section follows a methodical approach: gather information, analyse symptoms, identify root causes, and apply remediation.
flowchart TD
A[Issue Detected] --> B{Component Type?}
B -->|Pod| C[Pod Diagnostics]
B -->|Node| D[Node Health Checks]
B -->|Network| E[Network Debugging]
B -->|Storage| F[Storage Investigation]
B -->|Resources| G[Resource Analysis]
C --> H[Check Logs]
C --> I[Describe Pod]
C --> J[Check Events]
D --> K[Node Status]
D --> L[Taint Analysis]
D --> M[Drain/Cordon]
E --> N[DNS Testing]
E --> O[Connectivity Tests]
E --> P[Packet Capture]
F --> Q[PVC/PV Status]
F --> R[CSI Logs]
F --> S[Mount Issues]
G --> T[Limits/Requests]
G --> U[Eviction Events]
G --> V[Resource Quotas]
Pod Crash Diagnostics
Understanding Pod States
Pods transition through various states that indicate their health and availability. Common problematic states require different diagnostic approaches.
stateDiagram-v2
[*] --> Pending
Pending --> ContainerCreating : Image pulled
ContainerCreating --> Running : Container started
Running --> CrashLoopBackOff : Exit non-zero repeatedly
Running --> Error : Single failure
Pending --> ImagePullBackOff : Image pull failed
Pending --> ErrImagePull : Initial image error
ContainerCreating --> CreateContainerError : Container creation failed
Running --> OOMKilled : Out of memory
CrashLoopBackOff --> Running : Successful restart
ImagePullBackOff --> Running : Image available
Initial Diagnostics Workflow
# Get pod status with key information
kubectl get pods -o wide
# Check specific pod with custom columns
kubectl get pod <pod-name> -o custom-columns=\
NAME:.metadata.name,\
STATUS:.status.phase,\
RESTARTS:.status.containerStatuses[*].restartCount,\
NODE:.spec.nodeName
# View all pod conditions
kubectl get pod <pod-name> -o jsonpath='{.status.conditions[*].type}'
# Check container statuses within pod
kubectl get pod <pod-name> -o jsonpath='{.status.containerStatuses[*].state}'
Examining Pod Logs
Logs are the primary source of application-level diagnostic information.
# View current logs
kubectl logs <pod-name>
# View logs from specific container in multi-container pod
kubectl logs <pod-name> -c <container-name>
# View previous container logs (after crash)
kubectl logs <pod-name> --previous
# Follow logs in real-time
kubectl logs <pod-name> -f
# Get last N lines
kubectl logs <pod-name> --tail=100
# Logs with timestamps
kubectl logs <pod-name> --timestamps
# Logs since specific time
kubectl logs <pod-name> --since=1h
kubectl logs <pod-name> --since=2024-01-01T10:00:00Z
# All containers in pod
kubectl logs <pod-name> --all-containers=true
# Logs from all pods with specific label
kubectl logs -l app=myapp --all-containers=true
Detailed Pod Description
The describe command provides comprehensive information about pod configuration, status, and events.
# Full pod description
kubectl describe pod <pod-name>
# Focus on specific sections with grep
kubectl describe pod <pod-name> | grep -A 10 "Events:"
kubectl describe pod <pod-name> | grep -A 5 "Conditions:"
kubectl describe pod <pod-name> | grep -A 10 "Containers:"
# Check resource requests and limits
kubectl describe pod <pod-name> | grep -A 10 "Limits:"
# View container states
kubectl describe pod <pod-name> | grep -A 20 "State:"
Key sections to examine:
| Section | What to Check |
|---|---|
| Status | Current phase (Pending, Running, Failed) |
| IP | Pod has assigned IP address |
| Controlled By | Parent resource (Deployment, StatefulSet) |
| Containers | Image, ports, resource settings |
| Conditions | PodScheduled, Initialized, Ready, ContainersReady |
| State | Waiting, Running, Terminated with reason |
| Events | Recent operations and errors |
Analysing Events
Events provide a timeline of operations and errors affecting the pod.
# Get events for specific pod
kubectl get events --field-selector involvedObject.name=<pod-name>
# Events sorted by timestamp
kubectl get events --sort-by=.metadata.creationTimestamp
# Events from all namespaces
kubectl get events -A --sort-by=.metadata.creationTimestamp
# Warning events only
kubectl get events --field-selector type=Warning
# Events in last hour
kubectl get events --field-selector type=Warning \
--sort-by=.metadata.creationTimestamp | \
awk -v date="$(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%S)" '$1 > date'
# Watch events in real-time
kubectl get events -w
# Detailed event information
kubectl describe event <event-name>
Common Pod Failure Scenarios
CrashLoopBackOff
Container repeatedly crashes after starting.
# Check previous logs for crash reason
kubectl logs <pod-name> --previous
# Describe to see exit codes
kubectl describe pod <pod-name> | grep -A 10 "Last State:"
# Check for liveness probe failures
kubectl describe pod <pod-name> | grep -A 5 "Liveness:"
# Verify command and arguments
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].command}'
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].args}'
Common causes:
- Application startup errors
- Missing environment variables or secrets
- Failed liveness probe
- Insufficient permissions
- Missing dependencies
ImagePullBackOff
Cannot pull container image from registry.
# Check image name and tag
kubectl describe pod <pod-name> | grep -A 2 "Image:"
# Check pull policy
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].imagePullPolicy}'
# Verify image pull secrets
kubectl get pod <pod-name> -o jsonpath='{.spec.imagePullSecrets[*].name}'
# Inspect image pull secret
kubectl get secret <secret-name> -o yaml
# Test image pull manually on node
ssh <node>
docker pull <image-name> # or crictl pull
Common causes:
- Incorrect image name or tag
- Private registry without imagePullSecrets
- Network connectivity to registry
- Rate limiting from registry
- Authentication expired
Pending State
Pod cannot be scheduled onto a node.
# Check scheduling events
kubectl describe pod <pod-name> | grep -A 20 "Events:"
# View pod scheduling constraints
kubectl get pod <pod-name> -o yaml | grep -A 10 "nodeSelector:"
kubectl get pod <pod-name> -o yaml | grep -A 10 "affinity:"
kubectl get pod <pod-name> -o yaml | grep -A 10 "tolerations:"
# Check node resources
kubectl describe nodes | grep -A 5 "Allocated resources:"
# View taints on nodes
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
TAINTS:.spec.taints[*].effect
Common causes:
- Insufficient CPU or memory on nodes
- Node selector doesn't match any node
- Taints without corresponding tolerations
- Pod affinity/anti-affinity rules not satisfied
- PersistentVolumeClaim not bound
OOMKilled
Container terminated due to out-of-memory.
# Check container memory limits
kubectl describe pod <pod-name> | grep -A 5 "Limits:"
# View termination reason
kubectl describe pod <pod-name> | grep -A 10 "Last State:"
# Check node memory pressure
kubectl describe node <node-name> | grep "MemoryPressure"
# Review memory requests vs limits
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].resources}'
# Check for memory leaks in logs
kubectl logs <pod-name> --previous | grep -i "memory\|oom"
Remediation:
- Increase memory limits
- Investigate memory leaks in application
- Add memory requests for proper scheduling
- Monitor actual memory usage patterns
Interactive Debugging
# Execute command in running container
kubectl exec <pod-name> -- ls -la /app
# Interactive shell (if shell available)
kubectl exec -it <pod-name> -- /bin/bash
kubectl exec -it <pod-name> -- /bin/sh
# Multi-container pod - specify container
kubectl exec -it <pod-name> -c <container-name> -- /bin/bash
# Check process list
kubectl exec <pod-name> -- ps aux
# Verify environment variables
kubectl exec <pod-name> -- env
# Check network connectivity from pod
kubectl exec <pod-name> -- ping -c 3 google.com
kubectl exec <pod-name> -- nslookup kubernetes.default
# Inspect mounted volumes
kubectl exec <pod-name> -- df -h
kubectl exec <pod-name> -- ls -la /path/to/mount
Debug Containers (Kubernetes 1.23+)
Ephemeral containers for debugging pods without modifying the pod spec.
# Add debug container to running pod
kubectl debug <pod-name> -it --image=busybox --target=<container-name>
# Debug with different image
kubectl debug <pod-name> -it --image=ubuntu --share-processes
# Debug with network namespace sharing
kubectl debug <pod-name> -it --image=nicolaka/netshoot
# Create copy of pod for debugging (without sidecars)
kubectl debug <pod-name> -it --copy-to=<pod-name>-debug --container=debug
# Debug node by creating pod on specific node
kubectl debug node/<node-name> -it --image=ubuntu
Node Health Checks
Node Status Overview
# List all nodes with status
kubectl get nodes
# Detailed node information
kubectl get nodes -o wide
# Node status with conditions (filter by type; condition order is not guaranteed by the API)
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
READY:'.status.conditions[?(@.type=="Ready")].status',\
MEMORY:'.status.conditions[?(@.type=="MemoryPressure")].status',\
DISK:'.status.conditions[?(@.type=="DiskPressure")].status',\
PID:'.status.conditions[?(@.type=="PIDPressure")].status'
# Check node conditions
kubectl describe node <node-name> | grep -A 10 "Conditions:"
graph TB
A[Node Health Check] --> B{Node Ready?}
B -->|No| C[Check Conditions]
B -->|Yes| D[Check Resource Pressure]
C --> E{Memory Pressure?}
C --> F{Disk Pressure?}
C --> G{PID Pressure?}
C --> H{Network Unavailable?}
E -->|Yes| I[Free Memory / Evict Pods]
F -->|Yes| J[Clean Disk / Expand Storage]
G -->|Yes| K[Investigate High PID Usage]
H -->|Yes| L[Check Network Plugin]
D --> M[Review Allocatable Resources]
M --> N[Check Taints]
N --> O{Should Accept Pods?}
O -->|No| P[Cordon Node]
O -->|Yes| Q[Uncordon Node]
Node Conditions
Kubernetes tracks several node conditions:
| Condition | Meaning | Healthy Value |
|---|---|---|
| Ready | Node can accept pods | True |
| MemoryPressure | Node running out of memory | False |
| DiskPressure | Node running out of disk space | False |
| PIDPressure | Too many processes on node | False |
| NetworkUnavailable | Network not configured correctly | False |
# Check specific condition
kubectl get node <node-name> -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}'
# View all conditions with messages
kubectl describe node <node-name> | grep -A 15 "Conditions:"
# Check capacity and allocatable resources
kubectl describe node <node-name> | grep -A 10 "Capacity:"
kubectl describe node <node-name> | grep -A 10 "Allocatable:"
# View current resource allocation
kubectl describe node <node-name> | grep -A 15 "Allocated resources:"
Cordoning and Draining Nodes
Cordon prevents new pods from being scheduled on a node. Drain safely evicts existing pods.
# Cordon node (mark as unschedulable)
kubectl cordon <node-name>
# Verify node is cordoned
kubectl get nodes | grep SchedulingDisabled
# Uncordon node (mark as schedulable)
kubectl uncordon <node-name>
# Drain node (evict all pods)
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data
# Drain with grace period
kubectl drain <node-name> --grace-period=60 --ignore-daemonsets
# Drain including pods not managed by a controller
kubectl drain <node-name> --force --delete-emptydir-data
# Drain with pod deletion timeout
kubectl drain <node-name> --timeout=300s --ignore-daemonsets
# Dry run to see what would be evicted
kubectl drain <node-name> --dry-run=client --ignore-daemonsets
Drain workflow:
sequenceDiagram
participant A as Admin
participant N as Node
participant P as Pods
participant S as Scheduler
A->>N: kubectl cordon
N->>S: Mark Unschedulable
A->>N: kubectl drain
N->>P: Evict Pod 1
P->>P: Graceful Shutdown
P-->>N: Terminated
N->>P: Evict Pod 2
P->>P: Graceful Shutdown
P-->>N: Terminated
N-->>A: All Pods Evicted
A->>N: Perform Maintenance
A->>N: kubectl uncordon
N->>S: Mark Schedulable
Taint Analysis
Taints prevent pods from scheduling unless they have matching tolerations.
# View taints on all nodes
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
TAINTS:.spec.taints
# View taints in detail
kubectl describe node <node-name> | grep -A 5 "Taints:"
# Add taint to node
kubectl taint nodes <node-name> key=value:NoSchedule
# Taint effects
kubectl taint nodes <node-name> maintenance=true:NoSchedule # Don't schedule new
kubectl taint nodes <node-name> maintenance=true:PreferNoSchedule # Avoid if possible
kubectl taint nodes <node-name> maintenance=true:NoExecute # Evict existing
# Remove taint (note the minus sign)
kubectl taint nodes <node-name> key=value:NoSchedule-
# Remove all taints with specific key
kubectl taint nodes <node-name> key-
Common taints:
| Taint | Purpose |
|---|---|
node.kubernetes.io/not-ready |
Node not ready (automatic) |
node.kubernetes.io/unreachable |
Node unreachable (automatic) |
node.kubernetes.io/memory-pressure |
Low memory (automatic) |
node.kubernetes.io/disk-pressure |
Low disk space (automatic) |
node.kubernetes.io/network-unavailable |
Network issue (automatic) |
node.kubernetes.io/unschedulable |
Cordoned (automatic) |
Viewing Pod Tolerations
# Check pod tolerations
kubectl get pod <pod-name> -o jsonpath='{.spec.tolerations}'
# View tolerations in YAML
kubectl get pod <pod-name> -o yaml | grep -A 10 "tolerations:"
# Example toleration for tainted node
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: nginx-tolerant
spec:
tolerations:
- key: "maintenance"
operator: "Equal"
value: "true"
effect: "NoSchedule"
containers:
- name: nginx
image: nginx:1.25
EOF
Node Resource Monitoring
# Top nodes (requires metrics-server)
kubectl top nodes
# Top nodes with sorting
kubectl top nodes --sort-by=cpu
kubectl top nodes --sort-by=memory
# Detailed resource usage
kubectl describe node <node-name> | grep -A 20 "Allocated resources:"
# Calculate available resources
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
CPU_CAPACITY:.status.capacity.cpu,\
MEM_CAPACITY:.status.capacity.memory,\
CPU_ALLOCATABLE:.status.allocatable.cpu,\
MEM_ALLOCATABLE:.status.allocatable.memory
# View system pods consuming resources
kubectl top pods -n kube-system
# Check kubelet logs (on node)
journalctl -u kubelet -f
journalctl -u kubelet --since "1 hour ago"
Node Problem Detector
If Node Problem Detector is deployed, check for detected issues:
# Check node conditions added by NPD
kubectl describe node <node-name> | grep -i "kernel\|docker\|network"
# View NPD logs
kubectl logs -n kube-system -l app=node-problem-detector
# Check for specific problem conditions
kubectl get node <node-name> -o jsonpath='{.status.conditions[*].type}'
Network Debugging
DNS Troubleshooting
DNS resolution is critical for service discovery in Kubernetes.
flowchart LR
P[Pod] -->|1. Query| CD[CoreDNS/kube-dns]
CD -->|2. Check Cache| CD
CD -->|3. Query| US[Upstream Server]
US -->|4. Response| CD
CD -->|5. Return| P
style CD fill:#a8dadc
style P fill:#f1faee
style US fill:#e63946
# Test DNS from within a pod
kubectl run dnstest --image=busybox:1.28 --rm -it --restart=Never -- nslookup kubernetes.default
# Test service DNS resolution
kubectl run dnstest --image=busybox:1.28 --rm -it --restart=Never -- nslookup <service-name>.<namespace>.svc.cluster.local
# Check CoreDNS pods
kubectl get pods -n kube-system -l k8s-app=kube-dns
# View CoreDNS logs
kubectl logs -n kube-system -l k8s-app=kube-dns
# Check CoreDNS ConfigMap
kubectl get configmap coredns -n kube-system -o yaml
# Verify DNS service endpoint
kubectl get svc -n kube-system kube-dns
kubectl get endpoints -n kube-system kube-dns
# Check pod DNS configuration
kubectl exec <pod-name> -- cat /etc/resolv.conf
# Test external DNS resolution
kubectl exec <pod-name> -- nslookup google.com
# Check DNS policy for pod
kubectl get pod <pod-name> -o jsonpath='{.spec.dnsPolicy}'
Common DNS issues:
| Issue | Diagnostic | Fix |
|---|---|---|
| Service not resolving | Check service exists | Create/fix service |
| External names failing | Check CoreDNS logs | Verify upstream DNS |
| Timeout on queries | Check CoreDNS CPU/memory | Scale CoreDNS |
| Wrong IP returned | Check endpoints | Fix service selector |
Network Connectivity Testing
# Create network troubleshooting pod
kubectl run netshoot --image=nicolaka/netshoot -it --rm --restart=Never -- /bin/bash
# Or use netshoot as debug container
kubectl debug <pod-name> -it --image=nicolaka/netshoot
# Test connectivity to service
kubectl exec <pod-name> -- curl -v http://<service-name>:<port>
# Test TCP connectivity
kubectl exec <pod-name> -- nc -zv <service-name> <port>
# Test with timeout
kubectl exec <pod-name> -- timeout 5 curl http://<service-name>
# Check pod-to-pod connectivity
kubectl exec <source-pod> -- ping <destination-pod-ip>
# Test service ClusterIP
kubectl exec <pod-name> -- curl <service-cluster-ip>:<port>
# Verify routing
kubectl exec <pod-name> -- ip route
# Check network interfaces
kubectl exec <pod-name> -- ip addr
# Test external connectivity
kubectl exec <pod-name> -- curl -I https://www.google.com
Service and Endpoint Verification
# List services
kubectl get svc
# Describe service
kubectl describe svc <service-name>
# Check service endpoints
kubectl get endpoints <service-name>
# Verify endpoint IPs match pod IPs
kubectl get pods -o wide -l <label-selector>
kubectl get endpoints <service-name>
# Test service proxy
kubectl proxy --port=8080
curl http://localhost:8080/api/v1/namespaces/<namespace>/services/<service-name>/proxy/
# Check service selector matches pods
kubectl get svc <service-name> -o jsonpath='{.spec.selector}'
kubectl get pods --show-labels | grep <label>
Network Policy Debugging
# List network policies
kubectl get networkpolicies
# Describe network policy
kubectl describe networkpolicy <policy-name>
# Check if network policy affects pod
kubectl get pods --show-labels
kubectl describe networkpolicy <policy-name> | grep -A 5 "podSelector:"
# Verify CNI plugin logs
kubectl logs -n kube-system -l k8s-app=calico-node # Calico
kubectl logs -n kube-system -l app=flannel # Flannel
kubectl logs -n kube-system -l k8s-app=cilium # Cilium
# Test with policy temporarily removed
kubectl delete networkpolicy <policy-name>
# Test connectivity
kubectl apply -f <policy-file>.yaml # Restore
Packet Capture
Capture network traffic for detailed analysis.
# Using tcpdump in debug container
kubectl debug <pod-name> -it --image=nicolaka/netshoot -- tcpdump -i any -w /tmp/capture.pcap
# Capture specific port
kubectl debug <pod-name> -it --image=nicolaka/netshoot -- tcpdump -i any port 80
# Capture and display
kubectl debug <pod-name> -it --image=nicolaka/netshoot -- tcpdump -i any -n -XX
# On node - capture pod traffic (find pod's network namespace)
# SSH to node
nsenter -t <pid> -n tcpdump -i eth0
# Capture CNI traffic on node
tcpdump -i cni0 -n
# Save and download capture
kubectl cp <namespace>/<pod-name>:/tmp/capture.pcap ./capture.pcap
Ingress Troubleshooting
# Check ingress resources
kubectl get ingress
# Describe ingress
kubectl describe ingress <ingress-name>
# Check ingress controller logs
kubectl logs -n ingress-nginx -l app.kubernetes.io/component=controller
# Verify ingress class
kubectl get ingressclass
# Test ingress from outside cluster
curl -H "Host: <hostname>" http://<ingress-controller-ip>
# Check backend service
kubectl get ingress <ingress-name> -o jsonpath='{.spec.rules[*].http.paths[*].backend}'
# Verify TLS certificate (if HTTPS)
kubectl get secret <tls-secret-name> -o yaml
CNI Plugin Diagnostics
# Check CNI plugin pods
kubectl get pods -n kube-system -l <cni-label>
# Calico diagnostics
kubectl get pods -n kube-system -l k8s-app=calico-node
kubectl exec -n kube-system <calico-pod> -- calicoctl node status
kubectl exec -n kube-system <calico-pod> -- calicoctl get workloadEndpoint
# Cilium diagnostics
kubectl -n kube-system exec -it <cilium-pod> -- cilium status
kubectl -n kube-system exec -it <cilium-pod> -- cilium endpoint list
# Flannel diagnostics
kubectl logs -n kube-system -l app=flannel
# Check pod network CIDR configuration
kubectl cluster-info dump | grep -i cidr
Storage Issues
PersistentVolume and PersistentVolumeClaim States
stateDiagram-v2
[*] --> PVC_Pending : PVC Created
[*] --> PV_Available : PV Created
PV_Available --> PV_Bound : Claim Matches
PVC_Pending --> PVC_Bound : PV Bound
PV_Bound --> PV_Released : PVC Deleted
PV_Released --> PV_Available : Reclaimed
PV_Released --> PV_Failed : Reclaim Error
PVC_Bound --> [*] : PVC Deleted
PV_Available --> [*] : PV Deleted
PV_Failed --> [*]
Checking PVC and PV Status
# List PersistentVolumeClaims
kubectl get pvc
# Detailed PVC information
kubectl get pvc -o wide
# Describe PVC
kubectl describe pvc <pvc-name>
# Check PVC events
kubectl get events --field-selector involvedObject.name=<pvc-name>
# List PersistentVolumes
kubectl get pv
# Describe PV
kubectl describe pv <pv-name>
# Check PV capacity and status
kubectl get pv -o custom-columns=\
NAME:.metadata.name,\
CAPACITY:.spec.capacity.storage,\
STATUS:.status.phase,\
CLAIM:.spec.claimRef.name
# Verify PVC binding
kubectl get pvc <pvc-name> -o jsonpath='{.spec.volumeName}'
kubectl get pv <pv-name> -o jsonpath='{.spec.claimRef.name}'
PVC Pending State
Common reasons for PVC to remain in Pending state:
# Check if matching PV exists
kubectl get pv | grep Available
# Verify storage class exists
kubectl get storageclass
kubectl describe storageclass <storage-class-name>
# Check PVC requested capacity
kubectl get pvc <pvc-name> -o jsonpath='{.spec.resources.requests.storage}'
# Verify access modes match
kubectl get pvc <pvc-name> -o jsonpath='{.spec.accessModes}'
kubectl get pv -o custom-columns=NAME:.metadata.name,ACCESS:.spec.accessModes
# Check storage provisioner logs
kubectl logs -n kube-system -l app=<provisioner-name>
# Review PVC events for errors
kubectl describe pvc <pvc-name> | grep -A 10 "Events:"
Common PVC issues:
| Issue | Cause | Resolution |
|---|---|---|
| Pending | No matching PV | Create PV or use dynamic provisioning |
| Pending | StorageClass not found | Create StorageClass or fix name |
| Pending | Insufficient capacity | Create larger PV or resize |
| Pending | Access mode mismatch | Align PV and PVC access modes |
| Lost | PV deleted | Cannot recover; restore from backup |
Dynamic Provisioning Issues
# Check StorageClass configuration
kubectl get storageclass <sc-name> -o yaml
# Verify default StorageClass
kubectl get storageclass | grep default
# Set default StorageClass
kubectl patch storageclass <sc-name> -p \
'{"metadata": {"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'
# Check provisioner
kubectl get storageclass <sc-name> -o jsonpath='{.provisioner}'
# View volume binding mode
kubectl get storageclass <sc-name> -o jsonpath='{.volumeBindingMode}'
# WaitForFirstConsumer vs Immediate
# WaitForFirstConsumer: Volume created when pod is scheduled
# Immediate: Volume created when PVC is created
CSI Driver Diagnostics
Container Storage Interface (CSI) drivers handle volume operations.
# List CSI drivers
kubectl get csidrivers
# Check CSI controller pods
kubectl get pods -n kube-system -l app=<csi-driver>
# CSI node driver pods (on each node)
kubectl get pods -n kube-system -l app=<csi-driver-node>
# View CSI driver logs - controller
kubectl logs -n kube-system -l app=<csi-driver> -c <controller-container>
# View CSI driver logs - node
kubectl logs -n kube-system -l app=<csi-driver-node> -c <node-container>
# Check CSINode objects
kubectl get csinodes
# Describe CSINode for specific node
kubectl describe csinode <node-name>
# Check VolumeAttachments
kubectl get volumeattachments
# Describe VolumeAttachment
kubectl describe volumeattachment <va-name>
Common CSI issues:
# FailedAttachVolume
kubectl describe pod <pod-name> | grep FailedAttachVolume
# Check: VolumeAttachment status, CSI driver logs, node taints
# FailedMount
kubectl describe pod <pod-name> | grep FailedMount
# Check: Mount permissions, CSI node driver, volume exists
# Volume already attached
kubectl get volumeattachments | grep <pv-name>
# May need to manually detach or wait for cleanup
# CSI driver timeout
kubectl logs -n kube-system <csi-pod> | grep -i timeout
# Check: Network latency, storage backend health
Volume Mount Issues
# Check pod volume mounts
kubectl describe pod <pod-name> | grep -A 10 "Mounts:"
# Verify volumes defined in pod
kubectl describe pod <pod-name> | grep -A 10 "Volumes:"
# Check if volume is mounted in container
kubectl exec <pod-name> -- df -h
# Check mount permissions
kubectl exec <pod-name> -- ls -la /path/to/mount
# View mount details
kubectl exec <pod-name> -- mount | grep /path/to/mount
# On node - check mounted volumes
# SSH to node
mount | grep kubernetes
lsblk
Storage Capacity and Usage
# Check PV capacity
kubectl get pv -o custom-columns=NAME:.metadata.name,CAPACITY:.spec.capacity.storage
# View actual disk usage (requires metrics)
kubectl exec <pod-name> -- df -h /path/to/volume
# Check storage quotas in namespace
kubectl get resourcequota
# Describe resource quota
kubectl describe resourcequota <quota-name>
# View storage usage by PVC (if supported)
kubectl get pvc -o custom-columns=\
NAME:.metadata.name,\
CAPACITY:.status.capacity.storage,\
USED:.status.used
# Check node disk pressure
kubectl describe nodes | grep -i "disk"
Resizing Volumes
# Check if StorageClass allows expansion
kubectl get storageclass <sc-name> -o jsonpath='{.allowVolumeExpansion}'
# Expand PVC
kubectl patch pvc <pvc-name> -p \
'{"spec":{"resources":{"requests":{"storage":"20Gi"}}}}'
# Monitor resize operation
kubectl describe pvc <pvc-name> | grep -A 5 "Conditions:"
# Check for FileSystemResizePending condition
kubectl get pvc <pvc-name> -o jsonpath='{.status.conditions[?(@.type=="FileSystemResizePending")]}'
# Trigger filesystem resize (may require pod restart)
kubectl delete pod <pod-name>
# Verify new size
kubectl exec <pod-name> -- df -h /path/to/mount
Resource Pressure Remediation
Understanding Resource Requests and Limits
graph TB
subgraph "Resource Allocation"
A[No Request/Limit] -->|BestEffort| B[Evicted First]
C[Request Only] -->|Burstable| D[Evicted Second]
E[Request = Limit] -->|Guaranteed| F[Evicted Last]
end
subgraph "QoS Classes"
B
D
F
end
style B fill:#e63946
style D fill:#f4a261
style F fill:#2a9d8f
Viewing Resource Configuration
# Check pod resource requests and limits
kubectl describe pod <pod-name> | grep -A 10 "Requests:\|Limits:"
# Get resource settings in structured format
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].resources}'
# View QoS class
kubectl get pod <pod-name> -o jsonpath='{.status.qosClass}'
# All pods with QoS class
kubectl get pods -o custom-columns=\
NAME:.metadata.name,\
QOS:.status.qosClass,\
CPU_REQ:.spec.containers[*].resources.requests.cpu,\
MEM_REQ:.spec.containers[*].resources.requests.memory
# Check actual resource usage (requires metrics-server)
kubectl top pod <pod-name>
# Compare usage to limits
kubectl top pod <pod-name> --containers
QoS Classes
Kubernetes assigns QoS classes based on resource configuration:
| QoS Class | Criteria | Priority | Risk |
|---|---|---|---|
| Guaranteed | Requests = Limits for all resources | Highest | Lowest eviction risk |
| Burstable | At least one request/limit set | Medium | Medium eviction risk |
| BestEffort | No requests or limits | Lowest | Highest eviction risk |
# Find all BestEffort pods (high eviction risk)
kubectl get pods -o json | \
jq -r '.items[] | select(.status.qosClass=="BestEffort") | .metadata.name'
# Find all Burstable pods
kubectl get pods -o json | \
jq -r '.items[] | select(.status.qosClass=="Burstable") | .metadata.name'
# Find all Guaranteed pods
kubectl get pods -o json | \
jq -r '.items[] | select(.status.qosClass=="Guaranteed") | .metadata.name'
Node Resource Pressure
# Check node conditions for pressure
kubectl describe nodes | grep -E "MemoryPressure|DiskPressure|PIDPressure"
# Detailed node condition
kubectl get nodes -o jsonpath=\
'{range .items[*]}{.metadata.name}{"\t"}{.status.conditions[?(@.type=="MemoryPressure")].status}{"\n"}{end}'
# View node allocatable vs capacity
kubectl describe node <node-name> | grep -A 10 "Capacity:\|Allocatable:"
# Check current allocation percentage
kubectl describe node <node-name> | grep -A 15 "Allocated resources:"
# Top resource-consuming pods (kubectl top output has no node column;
# cross-reference with pods scheduled on the node)
kubectl top pods --all-namespaces --sort-by=memory
kubectl get pods --all-namespaces --field-selector spec.nodeName=<node-name>
Eviction Thresholds
Kubelet evicts pods when node resources are critically low.
# Check kubelet configuration for eviction thresholds (on node)
# Default thresholds:
# memory.available < 100Mi
# nodefs.available < 10%
# nodefs.inodesFree < 5%
# imagefs.available < 15%
# View kubelet config
kubectl proxy &
curl http://localhost:8001/api/v1/nodes/<node-name>/proxy/configz | jq .
# On node - check kubelet flags
ps aux | grep kubelet | grep eviction
# Or check kubelet config file
cat /var/lib/kubelet/config.yaml | grep -A 20 eviction
Diagnosing Evictions
# Find evicted pods
kubectl get pods --all-namespaces --field-selector=status.phase=Failed | grep Evicted
# Check eviction reason
kubectl describe pod <evicted-pod-name> | grep -A 5 "Reason:"
# Common reasons:
# - Evicted: Node was low on memory
# - Evicted: Node was low on disk space
# - Evicted: Node was low on ephemeral storage
# Get eviction events
kubectl get events --all-namespaces --field-selector reason=Evicted
# Detailed eviction information
kubectl get events --all-namespaces --field-selector reason=Evicted -o yaml
# Clean up evicted pods
kubectl delete pods --all-namespaces --field-selector=status.phase=Failed
Setting Appropriate Limits and Requests
# Analyse current usage to set appropriate values
kubectl top pod <pod-name> --containers
# Historical usage (requires monitoring system)
# Use Prometheus queries, Grafana dashboards, or cloud provider metrics
# Update deployment with resources
kubectl set resources deployment <deployment-name> \
--requests=cpu=100m,memory=128Mi \
--limits=cpu=500m,memory=512Mi
# Update specific container in multi-container pod
kubectl set resources deployment <deployment-name> \
-c <container-name> \
--requests=cpu=100m,memory=128Mi \
--limits=cpu=500m,memory=512Mi
# Patch pod resources
kubectl patch deployment <deployment-name> --patch '
spec:
template:
spec:
containers:
- name: <container-name>
resources:
requests:
memory: "128Mi"
cpu: "100m"
limits:
memory: "512Mi"
cpu: "500m"
'
Guidelines for setting resources:
- CPU requests: Base on average usage, allow bursting
- CPU limits: 2-4x requests, or omit for non-critical workloads
- Memory requests: Base on average usage + buffer
- Memory limits: Set to prevent OOM, typically 1.5-2x requests
- Start conservative: Monitor and adjust based on actual usage
LimitRanges
Enforce default limits and constraints in a namespace.
# View LimitRange in namespace
kubectl get limitrange
# Describe LimitRange
kubectl describe limitrange <limitrange-name>
# Example LimitRange
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: LimitRange
metadata:
name: default-limits
spec:
limits:
- default:
cpu: 500m
memory: 512Mi
defaultRequest:
cpu: 100m
memory: 128Mi
max:
cpu: "2"
memory: 2Gi
min:
cpu: 50m
memory: 64Mi
type: Container
EOF
# Check which pods are affected
kubectl get pods -o yaml | grep -A 10 "resources:"
ResourceQuotas
Limit total resource consumption in a namespace.
# View ResourceQuotas
kubectl get resourcequota
# Describe quota
kubectl describe resourcequota <quota-name>
# Check quota usage vs hard limits
kubectl get resourcequota <quota-name> -o jsonpath='{.status}'
# Example ResourceQuota
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
spec:
hard:
requests.cpu: "10"
requests.memory: 20Gi
limits.cpu: "20"
limits.memory: 40Gi
persistentvolumeclaims: "5"
pods: "20"
EOF
# Check if quota is preventing pod creation
kubectl describe resourcequota | grep -A 10 "Used:"
# Find which resources are at quota
kubectl get events --field-selector reason=FailedCreate | grep quota
Vertical Pod Autoscaling (VPA)
VPA recommends or automatically updates resource requests.
# Check if VPA is installed
kubectl get crd | grep verticalpodautoscaler
# List VPA objects
kubectl get vpa
# Describe VPA
kubectl describe vpa <vpa-name>
# View VPA recommendations
kubectl get vpa <vpa-name> -o jsonpath='{.status.recommendation}'
# Example VPA (recommendation mode)
cat <<EOF | kubectl apply -f -
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: nginx-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: nginx
updatePolicy:
updateMode: "Off" # Off, Initial, Recreate, Auto
EOF
# Check VPA recommendations
kubectl describe vpa nginx-vpa | grep -A 20 "Recommendation:"
Horizontal Pod Autoscaling (HPA)
HPA scales replica count based on metrics.
# List HPA objects
kubectl get hpa
# Describe HPA
kubectl describe hpa <hpa-name>
# Check HPA status and current metrics
kubectl get hpa <hpa-name> -o yaml
# Create HPA based on CPU
kubectl autoscale deployment <deployment-name> \
--cpu-percent=80 --min=2 --max=10
# Example HPA with multiple metrics
cat <<EOF | kubectl apply -f -
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: nginx-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: nginx
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 80
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
EOF
# Monitor HPA decisions
kubectl get hpa -w
# Check HPA events
kubectl describe hpa <hpa-name> | grep -A 10 "Events:"
Quick Reference
Common Diagnostic Commands
| Scenario | Command |
|---|---|
| Pod crashes | kubectl logs <pod> --previous |
| Pod pending | kubectl describe pod <pod> |
| Node issues | kubectl describe node <node> |
| DNS problems | kubectl logs -n kube-system -l k8s-app=kube-dns |
| Network test | kubectl run netshoot --image=nicolaka/netshoot --rm -it |
| PVC pending | kubectl describe pvc <pvc> |
| Resource usage | kubectl top pods --containers |
| Eviction events | kubectl get events --field-selector reason=Evicted |
| Debug container | kubectl debug <pod> -it --image=busybox |
| Service endpoints | kubectl get endpoints <service> |
Troubleshooting Decision Tree
flowchart TD
A[Problem Detected] --> B{Scope?}
B -->|Single Pod| C[Check Pod Status]
B -->|Multiple Pods| D[Check Node/Cluster]
B -->|Network| E[Test Connectivity]
B -->|Storage| F[Check PVC/PV]
C --> C1{State?}
C1 -->|CrashLoopBackOff| C2[Logs + Describe]
C1 -->|ImagePullBackOff| C3[Image + Secrets]
C1 -->|Pending| C4[Resources + Taints]
C1 -->|OOMKilled| C5[Memory Limits]
D --> D1[Node Conditions]
D1 --> D2[Resource Pressure?]
D2 -->|Yes| D3[Eviction/Limits]
D2 -->|No| D4[Cordon/Drain]
E --> E1[DNS Test]
E1 --> E2[Service/Endpoints]
E2 --> E3[Network Policy]
F --> F1[PVC Status]
F1 --> F2[Storage Class]
F2 --> F3[CSI Logs]
Resource Management Cheatsheet
| Task | Command |
|---|---|
| Set requests/limits | kubectl set resources deployment <name> --requests=cpu=100m,memory=128Mi --limits=cpu=500m,memory=512Mi |
| View QoS class | kubectl get pod <pod> -o jsonpath='{.status.qosClass}' |
| Top pods by memory | kubectl top pods --sort-by=memory |
| Top pods by CPU | kubectl top pods --sort-by=cpu |
| Create HPA | kubectl autoscale deployment <name> --cpu-percent=80 --min=2 --max=10 |
| View LimitRange | kubectl describe limitrange |
| View ResourceQuota | kubectl describe resourcequota |
| Find evicted pods | kubectl get pods -A --field-selector=status.phase=Failed | grep Evicted |
| Delete evicted pods | kubectl delete pods -A --field-selector=status.phase=Failed |
Node Management Cheatsheet
| Task | Command |
|---|---|
| Cordon node | kubectl cordon <node> |
| Uncordon node | kubectl uncordon <node> |
| Drain node | kubectl drain <node> --ignore-daemonsets --delete-emptydir-data |
| Add taint | kubectl taint nodes <node> key=value:NoSchedule |
| Remove taint | kubectl taint nodes <node> key=value:NoSchedule- |
| View node conditions | kubectl describe node <node> | grep Conditions: -A 10 |
| Top nodes | kubectl top nodes --sort-by=memory |
| Node events | kubectl get events --field-selector involvedObject.kind=Node |
Common Issues and Solutions
Issue: Pods Stuck in ImagePullBackOff
Symptoms:
- Pod status shows
ImagePullBackOfforErrImagePull - Events show "Failed to pull image"
Diagnosis:
kubectl describe pod <pod-name> | grep -A 5 "Events:"
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].image}'
Solutions:
- Verify image name and tag are correct
- Check imagePullSecrets are configured for private registries
- Validate registry credentials
- Test image pull on node manually
- Check network connectivity to registry
Issue: CrashLoopBackOff
Symptoms:
- Pod repeatedly restarts
- Status shows
CrashLoopBackOff - Restart count increases
Diagnosis:
kubectl logs <pod-name> --previous
kubectl describe pod <pod-name>
Solutions:
- Check application logs for errors
- Verify environment variables and secrets
- Review liveness probe configuration
- Check file system permissions
- Validate application dependencies
Issue: Service Not Reachable
Symptoms:
- Cannot connect to service
- Timeout errors from clients
- Service returns no response
Diagnosis:
kubectl get svc <service-name>
kubectl get endpoints <service-name>
kubectl describe svc <service-name>
Solutions:
- Verify service selector matches pod labels
- Check endpoints exist (pods are ready)
- Test connectivity from within cluster
- Verify port configuration
- Check NetworkPolicy rules
Issue: PVC Stuck in Pending
Symptoms:
- PVC remains in
Pendingstate - Pod cannot start due to volume mount failure
Diagnosis:
kubectl describe pvc <pvc-name>
kubectl get pv
kubectl get storageclass
Solutions:
- Create matching PersistentVolume
- Verify StorageClass exists and is correct
- Check storage provisioner is running
- Validate access modes match
- Ensure sufficient capacity available
Issue: Node NotReady
Symptoms:
- Node status shows
NotReady - Pods not scheduling on node
- Node conditions show pressure
Diagnosis:
kubectl describe node <node-name>
kubectl get node <node-name> -o jsonpath='{.status.conditions}'
journalctl -u kubelet # on node
Solutions:
- Check kubelet is running on node
- Verify network connectivity
- Clear disk space if DiskPressure
- Free memory if MemoryPressure
- Restart kubelet service
- Check CNI plugin status
Issue: DNS Resolution Failures
Symptoms:
- Services not resolving by name
- nslookup fails from pods
- Application cannot find dependencies
Diagnosis:
kubectl run dnstest --image=busybox:1.28 --rm -it -- nslookup kubernetes.default
kubectl logs -n kube-system -l k8s-app=kube-dns
kubectl get svc -n kube-system kube-dns
Solutions:
- Verify CoreDNS pods are running
- Check CoreDNS ConfigMap
- Validate DNS service endpoint
- Test upstream DNS resolution
- Scale CoreDNS if needed
Issue: High Resource Consumption
Symptoms:
- Pods being evicted
- Node showing resource pressure
- Slow cluster performance
Diagnosis:
kubectl top nodes
kubectl top pods --all-namespaces --sort-by=memory
kubectl describe node <node-name> | grep -A 20 "Allocated resources:"
Solutions:
- Set appropriate resource limits
- Implement HPA for auto-scaling
- Add more cluster capacity
- Identify and fix resource leaks
- Implement ResourceQuotas
- Use VPA for right-sizing
This comprehensive troubleshooting playbook provides systematic approaches to diagnose and resolve the most common Kubernetes issues. Regular practice with these diagnostic commands builds troubleshooting proficiency and reduces mean time to resolution (MTTR) in production environments.