Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Kubernetes Troubleshooting Playbook

Systematic approaches for diagnosing and resolving common Kubernetes cluster issues.

Kubernetes Troubleshooting Playbook

Systematic approaches for diagnosing and resolving common Kubernetes cluster issues.

Overview

Kubernetes troubleshooting requires understanding the interaction between control plane components, worker nodes, networking layers, and storage systems. This playbook provides structured diagnostic workflows for identifying and resolving issues across pods, nodes, networking, storage, and resource management. Each section follows a methodical approach: gather information, analyse symptoms, identify root causes, and apply remediation.

PodNodeNetworkStorageResourcesIssue DetectedComponent Type?Pod DiagnosticsNode Health ChecksNetwork DebuggingStorageInvestigationResource AnalysisCheck LogsDescribe PodCheck EventsNode StatusTaint AnalysisDrain/CordonDNS TestingConnectivity TestsPacket CapturePVC/PV StatusCSI LogsMount IssuesLimits/RequestsEviction EventsResource QuotasPodNodeNetworkStorageResourcesIssue DetectedComponent Type?Pod DiagnosticsNode Health ChecksNetwork DebuggingStorageInvestigationResource AnalysisCheck LogsDescribe PodCheck EventsNode StatusTaint AnalysisDrain/CordonDNS TestingConnectivity TestsPacket CapturePVC/PV StatusCSI LogsMount IssuesLimits/RequestsEviction EventsResource Quotas

Pod Crash Diagnostics

Understanding Pod States

Pods transition through various states that indicate their health and availability. Common problematic states require different diagnostic approaches.

Image pulledContainer startedExit non-zero repeatedlySingle failureImage pull failedInitial image errorContainer creation failedOut of memorySuccessful restartImage availablePendingContainerCreatingRunningCrashLoopBackOffErrorImagePullBackOffErrImagePullCreateContainerErrorOOMKilledImage pulledContainer startedExit non-zero repeatedlySingle failureImage pull failedInitial image errorContainer creation failedOut of memorySuccessful restartImage availablePendingContainerCreatingRunningCrashLoopBackOffErrorImagePullBackOffErrImagePullCreateContainerErrorOOMKilled

Initial Diagnostics Workflow

# Get pod status with key information
kubectl get pods -o wide

# Check specific pod with custom columns
kubectl get pod <pod-name> -o custom-columns=\
NAME:.metadata.name,\
STATUS:.status.phase,\
RESTARTS:.status.containerStatuses[*].restartCount,\
NODE:.spec.nodeName

# View all pod conditions
kubectl get pod <pod-name> -o jsonpath='{.status.conditions[*].type}'

# Check container statuses within pod
kubectl get pod <pod-name> -o jsonpath='{.status.containerStatuses[*].state}'

Examining Pod Logs

Logs are the primary source of application-level diagnostic information.

# View current logs
kubectl logs <pod-name>

# View logs from specific container in multi-container pod
kubectl logs <pod-name> -c <container-name>

# View previous container logs (after crash)
kubectl logs <pod-name> --previous

# Follow logs in real-time
kubectl logs <pod-name> -f

# Get last N lines
kubectl logs <pod-name> --tail=100

# Logs with timestamps
kubectl logs <pod-name> --timestamps

# Logs since specific time
kubectl logs <pod-name> --since=1h
kubectl logs <pod-name> --since=2024-01-01T10:00:00Z

# All containers in pod
kubectl logs <pod-name> --all-containers=true

# Logs from all pods with specific label
kubectl logs -l app=myapp --all-containers=true

Detailed Pod Description

The describe command provides comprehensive information about pod configuration, status, and events.

# Full pod description
kubectl describe pod <pod-name>

# Focus on specific sections with grep
kubectl describe pod <pod-name> | grep -A 10 "Events:"
kubectl describe pod <pod-name> | grep -A 5 "Conditions:"
kubectl describe pod <pod-name> | grep -A 10 "Containers:"

# Check resource requests and limits
kubectl describe pod <pod-name> | grep -A 10 "Limits:"

# View container states
kubectl describe pod <pod-name> | grep -A 20 "State:"

Key sections to examine:

Section What to Check
Status Current phase (Pending, Running, Failed)
IP Pod has assigned IP address
Controlled By Parent resource (Deployment, StatefulSet)
Containers Image, ports, resource settings
Conditions PodScheduled, Initialized, Ready, ContainersReady
State Waiting, Running, Terminated with reason
Events Recent operations and errors

Analysing Events

Events provide a timeline of operations and errors affecting the pod.

# Get events for specific pod
kubectl get events --field-selector involvedObject.name=<pod-name>

# Events sorted by timestamp
kubectl get events --sort-by=.metadata.creationTimestamp

# Events from all namespaces
kubectl get events -A --sort-by=.metadata.creationTimestamp

# Warning events only
kubectl get events --field-selector type=Warning

# Events in last hour
kubectl get events --field-selector type=Warning \
  --sort-by=.metadata.creationTimestamp | \
  awk -v date="$(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%S)" '$1 > date'

# Watch events in real-time
kubectl get events -w

# Detailed event information
kubectl describe event <event-name>

Common Pod Failure Scenarios

CrashLoopBackOff

Container repeatedly crashes after starting.

# Check previous logs for crash reason
kubectl logs <pod-name> --previous

# Describe to see exit codes
kubectl describe pod <pod-name> | grep -A 10 "Last State:"

# Check for liveness probe failures
kubectl describe pod <pod-name> | grep -A 5 "Liveness:"

# Verify command and arguments
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].command}'
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].args}'

Common causes:

  • Application startup errors
  • Missing environment variables or secrets
  • Failed liveness probe
  • Insufficient permissions
  • Missing dependencies

ImagePullBackOff

Cannot pull container image from registry.

# Check image name and tag
kubectl describe pod <pod-name> | grep -A 2 "Image:"

# Check pull policy
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].imagePullPolicy}'

# Verify image pull secrets
kubectl get pod <pod-name> -o jsonpath='{.spec.imagePullSecrets[*].name}'

# Inspect image pull secret
kubectl get secret <secret-name> -o yaml

# Test image pull manually on node
ssh <node>
docker pull <image-name>  # or crictl pull

Common causes:

  • Incorrect image name or tag
  • Private registry without imagePullSecrets
  • Network connectivity to registry
  • Rate limiting from registry
  • Authentication expired

Pending State

Pod cannot be scheduled onto a node.

# Check scheduling events
kubectl describe pod <pod-name> | grep -A 20 "Events:"

# View pod scheduling constraints
kubectl get pod <pod-name> -o yaml | grep -A 10 "nodeSelector:"
kubectl get pod <pod-name> -o yaml | grep -A 10 "affinity:"
kubectl get pod <pod-name> -o yaml | grep -A 10 "tolerations:"

# Check node resources
kubectl describe nodes | grep -A 5 "Allocated resources:"

# View taints on nodes
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
TAINTS:.spec.taints[*].effect

Common causes:

  • Insufficient CPU or memory on nodes
  • Node selector doesn't match any node
  • Taints without corresponding tolerations
  • Pod affinity/anti-affinity rules not satisfied
  • PersistentVolumeClaim not bound

OOMKilled

Container terminated due to out-of-memory.

# Check container memory limits
kubectl describe pod <pod-name> | grep -A 5 "Limits:"

# View termination reason
kubectl describe pod <pod-name> | grep -A 10 "Last State:"

# Check node memory pressure
kubectl describe node <node-name> | grep "MemoryPressure"

# Review memory requests vs limits
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].resources}'

# Check for memory leaks in logs
kubectl logs <pod-name> --previous | grep -i "memory\|oom"

Remediation:

  • Increase memory limits
  • Investigate memory leaks in application
  • Add memory requests for proper scheduling
  • Monitor actual memory usage patterns

Interactive Debugging

# Execute command in running container
kubectl exec <pod-name> -- ls -la /app

# Interactive shell (if shell available)
kubectl exec -it <pod-name> -- /bin/bash
kubectl exec -it <pod-name> -- /bin/sh

# Multi-container pod - specify container
kubectl exec -it <pod-name> -c <container-name> -- /bin/bash

# Check process list
kubectl exec <pod-name> -- ps aux

# Verify environment variables
kubectl exec <pod-name> -- env

# Check network connectivity from pod
kubectl exec <pod-name> -- ping -c 3 google.com
kubectl exec <pod-name> -- nslookup kubernetes.default

# Inspect mounted volumes
kubectl exec <pod-name> -- df -h
kubectl exec <pod-name> -- ls -la /path/to/mount

Debug Containers (Kubernetes 1.23+)

Ephemeral containers for debugging pods without modifying the pod spec.

# Add debug container to running pod
kubectl debug <pod-name> -it --image=busybox --target=<container-name>

# Debug with different image
kubectl debug <pod-name> -it --image=ubuntu --share-processes

# Debug with network namespace sharing
kubectl debug <pod-name> -it --image=nicolaka/netshoot

# Create copy of pod for debugging (without sidecars)
kubectl debug <pod-name> -it --copy-to=<pod-name>-debug --container=debug

# Debug node by creating pod on specific node
kubectl debug node/<node-name> -it --image=ubuntu

Node Health Checks

Node Status Overview

# List all nodes with status
kubectl get nodes

# Detailed node information
kubectl get nodes -o wide

# Node status with conditions (filter by type; condition order is not guaranteed by the API)
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
READY:'.status.conditions[?(@.type=="Ready")].status',\
MEMORY:'.status.conditions[?(@.type=="MemoryPressure")].status',\
DISK:'.status.conditions[?(@.type=="DiskPressure")].status',\
PID:'.status.conditions[?(@.type=="PIDPressure")].status'

# Check node conditions
kubectl describe node <node-name> | grep -A 10 "Conditions:"
NoYesYesYesYesYesNoYesNode Health CheckNode Ready?Check ConditionsCheck ResourcePressureMemory Pressure?Disk Pressure?PID Pressure?Network Unavailable?Free Memory / EvictPodsClean Disk / ExpandStorageInvestigate High PIDUsageCheck Network PluginReview AllocatableResourcesCheck TaintsShould Accept Pods?Cordon NodeUncordon NodeNoYesYesYesYesYesNoYesNode Health CheckNode Ready?Check ConditionsCheck ResourcePressureMemory Pressure?Disk Pressure?PID Pressure?Network Unavailable?Free Memory / EvictPodsClean Disk / ExpandStorageInvestigate High PIDUsageCheck Network PluginReview AllocatableResourcesCheck TaintsShould Accept Pods?Cordon NodeUncordon Node

Node Conditions

Kubernetes tracks several node conditions:

Condition Meaning Healthy Value
Ready Node can accept pods True
MemoryPressure Node running out of memory False
DiskPressure Node running out of disk space False
PIDPressure Too many processes on node False
NetworkUnavailable Network not configured correctly False
# Check specific condition
kubectl get node <node-name> -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}'

# View all conditions with messages
kubectl describe node <node-name> | grep -A 15 "Conditions:"

# Check capacity and allocatable resources
kubectl describe node <node-name> | grep -A 10 "Capacity:"
kubectl describe node <node-name> | grep -A 10 "Allocatable:"

# View current resource allocation
kubectl describe node <node-name> | grep -A 15 "Allocated resources:"

Cordoning and Draining Nodes

Cordon prevents new pods from being scheduled on a node. Drain safely evicts existing pods.

# Cordon node (mark as unschedulable)
kubectl cordon <node-name>

# Verify node is cordoned
kubectl get nodes | grep SchedulingDisabled

# Uncordon node (mark as schedulable)
kubectl uncordon <node-name>

# Drain node (evict all pods)
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data

# Drain with grace period
kubectl drain <node-name> --grace-period=60 --ignore-daemonsets

# Drain including pods not managed by a controller
kubectl drain <node-name> --force --delete-emptydir-data

# Drain with pod deletion timeout
kubectl drain <node-name> --timeout=300s --ignore-daemonsets

# Dry run to see what would be evicted
kubectl drain <node-name> --dry-run=client --ignore-daemonsets

Drain workflow:

SchedulerPodsNodeAdminSchedulerPodsNodeAdminkubectl cordonMark Unschedulablekubectl drainEvict Pod 1Graceful ShutdownTerminatedEvict Pod 2Graceful ShutdownTerminatedAll Pods EvictedPerform Maintenancekubectl uncordonMark SchedulableSchedulerPodsNodeAdminSchedulerPodsNodeAdminkubectl cordonMark Unschedulablekubectl drainEvict Pod 1Graceful ShutdownTerminatedEvict Pod 2Graceful ShutdownTerminatedAll Pods EvictedPerform Maintenancekubectl uncordonMark Schedulable

Taint Analysis

Taints prevent pods from scheduling unless they have matching tolerations.

# View taints on all nodes
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
TAINTS:.spec.taints

# View taints in detail
kubectl describe node <node-name> | grep -A 5 "Taints:"

# Add taint to node
kubectl taint nodes <node-name> key=value:NoSchedule

# Taint effects
kubectl taint nodes <node-name> maintenance=true:NoSchedule       # Don't schedule new
kubectl taint nodes <node-name> maintenance=true:PreferNoSchedule # Avoid if possible
kubectl taint nodes <node-name> maintenance=true:NoExecute        # Evict existing

# Remove taint (note the minus sign)
kubectl taint nodes <node-name> key=value:NoSchedule-

# Remove all taints with specific key
kubectl taint nodes <node-name> key-

Common taints:

Taint Purpose
node.kubernetes.io/not-ready Node not ready (automatic)
node.kubernetes.io/unreachable Node unreachable (automatic)
node.kubernetes.io/memory-pressure Low memory (automatic)
node.kubernetes.io/disk-pressure Low disk space (automatic)
node.kubernetes.io/network-unavailable Network issue (automatic)
node.kubernetes.io/unschedulable Cordoned (automatic)

Viewing Pod Tolerations

# Check pod tolerations
kubectl get pod <pod-name> -o jsonpath='{.spec.tolerations}'

# View tolerations in YAML
kubectl get pod <pod-name> -o yaml | grep -A 10 "tolerations:"

# Example toleration for tainted node
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
  name: nginx-tolerant
spec:
  tolerations:
  - key: "maintenance"
    operator: "Equal"
    value: "true"
    effect: "NoSchedule"
  containers:
  - name: nginx
    image: nginx:1.25
EOF

Node Resource Monitoring

# Top nodes (requires metrics-server)
kubectl top nodes

# Top nodes with sorting
kubectl top nodes --sort-by=cpu
kubectl top nodes --sort-by=memory

# Detailed resource usage
kubectl describe node <node-name> | grep -A 20 "Allocated resources:"

# Calculate available resources
kubectl get nodes -o custom-columns=\
NAME:.metadata.name,\
CPU_CAPACITY:.status.capacity.cpu,\
MEM_CAPACITY:.status.capacity.memory,\
CPU_ALLOCATABLE:.status.allocatable.cpu,\
MEM_ALLOCATABLE:.status.allocatable.memory

# View system pods consuming resources
kubectl top pods -n kube-system

# Check kubelet logs (on node)
journalctl -u kubelet -f
journalctl -u kubelet --since "1 hour ago"

Node Problem Detector

If Node Problem Detector is deployed, check for detected issues:

# Check node conditions added by NPD
kubectl describe node <node-name> | grep -i "kernel\|docker\|network"

# View NPD logs
kubectl logs -n kube-system -l app=node-problem-detector

# Check for specific problem conditions
kubectl get node <node-name> -o jsonpath='{.status.conditions[*].type}'

Network Debugging

DNS Troubleshooting

DNS resolution is critical for service discovery in Kubernetes.

1. Query2. Check Cache3. Query4. Response5. ReturnPodCoreDNS/kube-dnsUpstream Server1. Query2. Check Cache3. Query4. Response5. ReturnPodCoreDNS/kube-dnsUpstream Server
# Test DNS from within a pod
kubectl run dnstest --image=busybox:1.28 --rm -it --restart=Never -- nslookup kubernetes.default

# Test service DNS resolution
kubectl run dnstest --image=busybox:1.28 --rm -it --restart=Never -- nslookup <service-name>.<namespace>.svc.cluster.local

# Check CoreDNS pods
kubectl get pods -n kube-system -l k8s-app=kube-dns

# View CoreDNS logs
kubectl logs -n kube-system -l k8s-app=kube-dns

# Check CoreDNS ConfigMap
kubectl get configmap coredns -n kube-system -o yaml

# Verify DNS service endpoint
kubectl get svc -n kube-system kube-dns
kubectl get endpoints -n kube-system kube-dns

# Check pod DNS configuration
kubectl exec <pod-name> -- cat /etc/resolv.conf

# Test external DNS resolution
kubectl exec <pod-name> -- nslookup google.com

# Check DNS policy for pod
kubectl get pod <pod-name> -o jsonpath='{.spec.dnsPolicy}'

Common DNS issues:

Issue Diagnostic Fix
Service not resolving Check service exists Create/fix service
External names failing Check CoreDNS logs Verify upstream DNS
Timeout on queries Check CoreDNS CPU/memory Scale CoreDNS
Wrong IP returned Check endpoints Fix service selector

Network Connectivity Testing

# Create network troubleshooting pod
kubectl run netshoot --image=nicolaka/netshoot -it --rm --restart=Never -- /bin/bash

# Or use netshoot as debug container
kubectl debug <pod-name> -it --image=nicolaka/netshoot

# Test connectivity to service
kubectl exec <pod-name> -- curl -v http://<service-name>:<port>

# Test TCP connectivity
kubectl exec <pod-name> -- nc -zv <service-name> <port>

# Test with timeout
kubectl exec <pod-name> -- timeout 5 curl http://<service-name>

# Check pod-to-pod connectivity
kubectl exec <source-pod> -- ping <destination-pod-ip>

# Test service ClusterIP
kubectl exec <pod-name> -- curl <service-cluster-ip>:<port>

# Verify routing
kubectl exec <pod-name> -- ip route

# Check network interfaces
kubectl exec <pod-name> -- ip addr

# Test external connectivity
kubectl exec <pod-name> -- curl -I https://www.google.com

Service and Endpoint Verification

# List services
kubectl get svc

# Describe service
kubectl describe svc <service-name>

# Check service endpoints
kubectl get endpoints <service-name>

# Verify endpoint IPs match pod IPs
kubectl get pods -o wide -l <label-selector>
kubectl get endpoints <service-name>

# Test service proxy
kubectl proxy --port=8080
curl http://localhost:8080/api/v1/namespaces/<namespace>/services/<service-name>/proxy/

# Check service selector matches pods
kubectl get svc <service-name> -o jsonpath='{.spec.selector}'
kubectl get pods --show-labels | grep <label>

Network Policy Debugging

# List network policies
kubectl get networkpolicies

# Describe network policy
kubectl describe networkpolicy <policy-name>

# Check if network policy affects pod
kubectl get pods --show-labels
kubectl describe networkpolicy <policy-name> | grep -A 5 "podSelector:"

# Verify CNI plugin logs
kubectl logs -n kube-system -l k8s-app=calico-node   # Calico
kubectl logs -n kube-system -l app=flannel           # Flannel
kubectl logs -n kube-system -l k8s-app=cilium        # Cilium

# Test with policy temporarily removed
kubectl delete networkpolicy <policy-name>
# Test connectivity
kubectl apply -f <policy-file>.yaml  # Restore

Packet Capture

Capture network traffic for detailed analysis.

# Using tcpdump in debug container
kubectl debug <pod-name> -it --image=nicolaka/netshoot -- tcpdump -i any -w /tmp/capture.pcap

# Capture specific port
kubectl debug <pod-name> -it --image=nicolaka/netshoot -- tcpdump -i any port 80

# Capture and display
kubectl debug <pod-name> -it --image=nicolaka/netshoot -- tcpdump -i any -n -XX

# On node - capture pod traffic (find pod's network namespace)
# SSH to node
nsenter -t <pid> -n tcpdump -i eth0

# Capture CNI traffic on node
tcpdump -i cni0 -n

# Save and download capture
kubectl cp <namespace>/<pod-name>:/tmp/capture.pcap ./capture.pcap

Ingress Troubleshooting

# Check ingress resources
kubectl get ingress

# Describe ingress
kubectl describe ingress <ingress-name>

# Check ingress controller logs
kubectl logs -n ingress-nginx -l app.kubernetes.io/component=controller

# Verify ingress class
kubectl get ingressclass

# Test ingress from outside cluster
curl -H "Host: <hostname>" http://<ingress-controller-ip>

# Check backend service
kubectl get ingress <ingress-name> -o jsonpath='{.spec.rules[*].http.paths[*].backend}'

# Verify TLS certificate (if HTTPS)
kubectl get secret <tls-secret-name> -o yaml

CNI Plugin Diagnostics

# Check CNI plugin pods
kubectl get pods -n kube-system -l <cni-label>

# Calico diagnostics
kubectl get pods -n kube-system -l k8s-app=calico-node
kubectl exec -n kube-system <calico-pod> -- calicoctl node status
kubectl exec -n kube-system <calico-pod> -- calicoctl get workloadEndpoint

# Cilium diagnostics
kubectl -n kube-system exec -it <cilium-pod> -- cilium status
kubectl -n kube-system exec -it <cilium-pod> -- cilium endpoint list

# Flannel diagnostics
kubectl logs -n kube-system -l app=flannel

# Check pod network CIDR configuration
kubectl cluster-info dump | grep -i cidr

Storage Issues

PersistentVolume and PersistentVolumeClaim States

PVC CreatedPV CreatedClaim MatchesPV BoundPVC DeletedReclaimedReclaim ErrorPVC DeletedPV DeletedPVC_PendingPV_AvailablePV_BoundPVC_BoundPV_ReleasedPV_FailedPVC CreatedPV CreatedClaim MatchesPV BoundPVC DeletedReclaimedReclaim ErrorPVC DeletedPV DeletedPVC_PendingPV_AvailablePV_BoundPVC_BoundPV_ReleasedPV_Failed

Checking PVC and PV Status

# List PersistentVolumeClaims
kubectl get pvc

# Detailed PVC information
kubectl get pvc -o wide

# Describe PVC
kubectl describe pvc <pvc-name>

# Check PVC events
kubectl get events --field-selector involvedObject.name=<pvc-name>

# List PersistentVolumes
kubectl get pv

# Describe PV
kubectl describe pv <pv-name>

# Check PV capacity and status
kubectl get pv -o custom-columns=\
NAME:.metadata.name,\
CAPACITY:.spec.capacity.storage,\
STATUS:.status.phase,\
CLAIM:.spec.claimRef.name

# Verify PVC binding
kubectl get pvc <pvc-name> -o jsonpath='{.spec.volumeName}'
kubectl get pv <pv-name> -o jsonpath='{.spec.claimRef.name}'

PVC Pending State

Common reasons for PVC to remain in Pending state:

# Check if matching PV exists
kubectl get pv | grep Available

# Verify storage class exists
kubectl get storageclass
kubectl describe storageclass <storage-class-name>

# Check PVC requested capacity
kubectl get pvc <pvc-name> -o jsonpath='{.spec.resources.requests.storage}'

# Verify access modes match
kubectl get pvc <pvc-name> -o jsonpath='{.spec.accessModes}'
kubectl get pv -o custom-columns=NAME:.metadata.name,ACCESS:.spec.accessModes

# Check storage provisioner logs
kubectl logs -n kube-system -l app=<provisioner-name>

# Review PVC events for errors
kubectl describe pvc <pvc-name> | grep -A 10 "Events:"

Common PVC issues:

Issue Cause Resolution
Pending No matching PV Create PV or use dynamic provisioning
Pending StorageClass not found Create StorageClass or fix name
Pending Insufficient capacity Create larger PV or resize
Pending Access mode mismatch Align PV and PVC access modes
Lost PV deleted Cannot recover; restore from backup

Dynamic Provisioning Issues

# Check StorageClass configuration
kubectl get storageclass <sc-name> -o yaml

# Verify default StorageClass
kubectl get storageclass | grep default

# Set default StorageClass
kubectl patch storageclass <sc-name> -p \
  '{"metadata": {"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'

# Check provisioner
kubectl get storageclass <sc-name> -o jsonpath='{.provisioner}'

# View volume binding mode
kubectl get storageclass <sc-name> -o jsonpath='{.volumeBindingMode}'

# WaitForFirstConsumer vs Immediate
# WaitForFirstConsumer: Volume created when pod is scheduled
# Immediate: Volume created when PVC is created

CSI Driver Diagnostics

Container Storage Interface (CSI) drivers handle volume operations.

# List CSI drivers
kubectl get csidrivers

# Check CSI controller pods
kubectl get pods -n kube-system -l app=<csi-driver>

# CSI node driver pods (on each node)
kubectl get pods -n kube-system -l app=<csi-driver-node>

# View CSI driver logs - controller
kubectl logs -n kube-system -l app=<csi-driver> -c <controller-container>

# View CSI driver logs - node
kubectl logs -n kube-system -l app=<csi-driver-node> -c <node-container>

# Check CSINode objects
kubectl get csinodes

# Describe CSINode for specific node
kubectl describe csinode <node-name>

# Check VolumeAttachments
kubectl get volumeattachments

# Describe VolumeAttachment
kubectl describe volumeattachment <va-name>

Common CSI issues:

# FailedAttachVolume
kubectl describe pod <pod-name> | grep FailedAttachVolume
# Check: VolumeAttachment status, CSI driver logs, node taints

# FailedMount
kubectl describe pod <pod-name> | grep FailedMount
# Check: Mount permissions, CSI node driver, volume exists

# Volume already attached
kubectl get volumeattachments | grep <pv-name>
# May need to manually detach or wait for cleanup

# CSI driver timeout
kubectl logs -n kube-system <csi-pod> | grep -i timeout
# Check: Network latency, storage backend health

Volume Mount Issues

# Check pod volume mounts
kubectl describe pod <pod-name> | grep -A 10 "Mounts:"

# Verify volumes defined in pod
kubectl describe pod <pod-name> | grep -A 10 "Volumes:"

# Check if volume is mounted in container
kubectl exec <pod-name> -- df -h

# Check mount permissions
kubectl exec <pod-name> -- ls -la /path/to/mount

# View mount details
kubectl exec <pod-name> -- mount | grep /path/to/mount

# On node - check mounted volumes
# SSH to node
mount | grep kubernetes
lsblk

Storage Capacity and Usage

# Check PV capacity
kubectl get pv -o custom-columns=NAME:.metadata.name,CAPACITY:.spec.capacity.storage

# View actual disk usage (requires metrics)
kubectl exec <pod-name> -- df -h /path/to/volume

# Check storage quotas in namespace
kubectl get resourcequota

# Describe resource quota
kubectl describe resourcequota <quota-name>

# View storage usage by PVC (if supported)
kubectl get pvc -o custom-columns=\
NAME:.metadata.name,\
CAPACITY:.status.capacity.storage,\
USED:.status.used

# Check node disk pressure
kubectl describe nodes | grep -i "disk"

Resizing Volumes

# Check if StorageClass allows expansion
kubectl get storageclass <sc-name> -o jsonpath='{.allowVolumeExpansion}'

# Expand PVC
kubectl patch pvc <pvc-name> -p \
  '{"spec":{"resources":{"requests":{"storage":"20Gi"}}}}'

# Monitor resize operation
kubectl describe pvc <pvc-name> | grep -A 5 "Conditions:"

# Check for FileSystemResizePending condition
kubectl get pvc <pvc-name> -o jsonpath='{.status.conditions[?(@.type=="FileSystemResizePending")]}'

# Trigger filesystem resize (may require pod restart)
kubectl delete pod <pod-name>

# Verify new size
kubectl exec <pod-name> -- df -h /path/to/mount

Resource Pressure Remediation

Understanding Resource Requests and Limits

QoS ClassesResource AllocationBestEffortBurstableGuaranteedNo Request/LimitEvicted FirstRequest OnlyEvicted SecondRequest = LimitEvicted LastQoS ClassesResource AllocationBestEffortBurstableGuaranteedNo Request/LimitEvicted FirstRequest OnlyEvicted SecondRequest = LimitEvicted Last

Viewing Resource Configuration

# Check pod resource requests and limits
kubectl describe pod <pod-name> | grep -A 10 "Requests:\|Limits:"

# Get resource settings in structured format
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].resources}'

# View QoS class
kubectl get pod <pod-name> -o jsonpath='{.status.qosClass}'

# All pods with QoS class
kubectl get pods -o custom-columns=\
NAME:.metadata.name,\
QOS:.status.qosClass,\
CPU_REQ:.spec.containers[*].resources.requests.cpu,\
MEM_REQ:.spec.containers[*].resources.requests.memory

# Check actual resource usage (requires metrics-server)
kubectl top pod <pod-name>

# Compare usage to limits
kubectl top pod <pod-name> --containers

QoS Classes

Kubernetes assigns QoS classes based on resource configuration:

QoS Class Criteria Priority Risk
Guaranteed Requests = Limits for all resources Highest Lowest eviction risk
Burstable At least one request/limit set Medium Medium eviction risk
BestEffort No requests or limits Lowest Highest eviction risk
# Find all BestEffort pods (high eviction risk)
kubectl get pods -o json | \
  jq -r '.items[] | select(.status.qosClass=="BestEffort") | .metadata.name'

# Find all Burstable pods
kubectl get pods -o json | \
  jq -r '.items[] | select(.status.qosClass=="Burstable") | .metadata.name'

# Find all Guaranteed pods
kubectl get pods -o json | \
  jq -r '.items[] | select(.status.qosClass=="Guaranteed") | .metadata.name'

Node Resource Pressure

# Check node conditions for pressure
kubectl describe nodes | grep -E "MemoryPressure|DiskPressure|PIDPressure"

# Detailed node condition
kubectl get nodes -o jsonpath=\
'{range .items[*]}{.metadata.name}{"\t"}{.status.conditions[?(@.type=="MemoryPressure")].status}{"\n"}{end}'

# View node allocatable vs capacity
kubectl describe node <node-name> | grep -A 10 "Capacity:\|Allocatable:"

# Check current allocation percentage
kubectl describe node <node-name> | grep -A 15 "Allocated resources:"

# Top resource-consuming pods (kubectl top output has no node column;
# cross-reference with pods scheduled on the node)
kubectl top pods --all-namespaces --sort-by=memory
kubectl get pods --all-namespaces --field-selector spec.nodeName=<node-name>

Eviction Thresholds

Kubelet evicts pods when node resources are critically low.

# Check kubelet configuration for eviction thresholds (on node)
# Default thresholds:
# memory.available < 100Mi
# nodefs.available < 10%
# nodefs.inodesFree < 5%
# imagefs.available < 15%

# View kubelet config
kubectl proxy &
curl http://localhost:8001/api/v1/nodes/<node-name>/proxy/configz | jq .

# On node - check kubelet flags
ps aux | grep kubelet | grep eviction

# Or check kubelet config file
cat /var/lib/kubelet/config.yaml | grep -A 20 eviction

Diagnosing Evictions

# Find evicted pods
kubectl get pods --all-namespaces --field-selector=status.phase=Failed | grep Evicted

# Check eviction reason
kubectl describe pod <evicted-pod-name> | grep -A 5 "Reason:"

# Common reasons:
# - Evicted: Node was low on memory
# - Evicted: Node was low on disk space
# - Evicted: Node was low on ephemeral storage

# Get eviction events
kubectl get events --all-namespaces --field-selector reason=Evicted

# Detailed eviction information
kubectl get events --all-namespaces --field-selector reason=Evicted -o yaml

# Clean up evicted pods
kubectl delete pods --all-namespaces --field-selector=status.phase=Failed

Setting Appropriate Limits and Requests

# Analyse current usage to set appropriate values
kubectl top pod <pod-name> --containers

# Historical usage (requires monitoring system)
# Use Prometheus queries, Grafana dashboards, or cloud provider metrics

# Update deployment with resources
kubectl set resources deployment <deployment-name> \
  --requests=cpu=100m,memory=128Mi \
  --limits=cpu=500m,memory=512Mi

# Update specific container in multi-container pod
kubectl set resources deployment <deployment-name> \
  -c <container-name> \
  --requests=cpu=100m,memory=128Mi \
  --limits=cpu=500m,memory=512Mi

# Patch pod resources
kubectl patch deployment <deployment-name> --patch '
spec:
  template:
    spec:
      containers:
      - name: <container-name>
        resources:
          requests:
            memory: "128Mi"
            cpu: "100m"
          limits:
            memory: "512Mi"
            cpu: "500m"
'

Guidelines for setting resources:

  • CPU requests: Base on average usage, allow bursting
  • CPU limits: 2-4x requests, or omit for non-critical workloads
  • Memory requests: Base on average usage + buffer
  • Memory limits: Set to prevent OOM, typically 1.5-2x requests
  • Start conservative: Monitor and adjust based on actual usage

LimitRanges

Enforce default limits and constraints in a namespace.

# View LimitRange in namespace
kubectl get limitrange

# Describe LimitRange
kubectl describe limitrange <limitrange-name>

# Example LimitRange
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: LimitRange
metadata:
  name: default-limits
spec:
  limits:
  - default:
      cpu: 500m
      memory: 512Mi
    defaultRequest:
      cpu: 100m
      memory: 128Mi
    max:
      cpu: "2"
      memory: 2Gi
    min:
      cpu: 50m
      memory: 64Mi
    type: Container
EOF

# Check which pods are affected
kubectl get pods -o yaml | grep -A 10 "resources:"

ResourceQuotas

Limit total resource consumption in a namespace.

# View ResourceQuotas
kubectl get resourcequota

# Describe quota
kubectl describe resourcequota <quota-name>

# Check quota usage vs hard limits
kubectl get resourcequota <quota-name> -o jsonpath='{.status}'

# Example ResourceQuota
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: ResourceQuota
metadata:
  name: compute-quota
spec:
  hard:
    requests.cpu: "10"
    requests.memory: 20Gi
    limits.cpu: "20"
    limits.memory: 40Gi
    persistentvolumeclaims: "5"
    pods: "20"
EOF

# Check if quota is preventing pod creation
kubectl describe resourcequota | grep -A 10 "Used:"

# Find which resources are at quota
kubectl get events --field-selector reason=FailedCreate | grep quota

Vertical Pod Autoscaling (VPA)

VPA recommends or automatically updates resource requests.

# Check if VPA is installed
kubectl get crd | grep verticalpodautoscaler

# List VPA objects
kubectl get vpa

# Describe VPA
kubectl describe vpa <vpa-name>

# View VPA recommendations
kubectl get vpa <vpa-name> -o jsonpath='{.status.recommendation}'

# Example VPA (recommendation mode)
cat <<EOF | kubectl apply -f -
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: nginx-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: nginx
  updatePolicy:
    updateMode: "Off"  # Off, Initial, Recreate, Auto
EOF

# Check VPA recommendations
kubectl describe vpa nginx-vpa | grep -A 20 "Recommendation:"

Horizontal Pod Autoscaling (HPA)

HPA scales replica count based on metrics.

# List HPA objects
kubectl get hpa

# Describe HPA
kubectl describe hpa <hpa-name>

# Check HPA status and current metrics
kubectl get hpa <hpa-name> -o yaml

# Create HPA based on CPU
kubectl autoscale deployment <deployment-name> \
  --cpu-percent=80 --min=2 --max=10

# Example HPA with multiple metrics
cat <<EOF | kubectl apply -f -
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: nginx-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: nginx
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 80
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 80
EOF

# Monitor HPA decisions
kubectl get hpa -w

# Check HPA events
kubectl describe hpa <hpa-name> | grep -A 10 "Events:"

Quick Reference

Common Diagnostic Commands

Scenario Command
Pod crashes kubectl logs <pod> --previous
Pod pending kubectl describe pod <pod>
Node issues kubectl describe node <node>
DNS problems kubectl logs -n kube-system -l k8s-app=kube-dns
Network test kubectl run netshoot --image=nicolaka/netshoot --rm -it
PVC pending kubectl describe pvc <pvc>
Resource usage kubectl top pods --containers
Eviction events kubectl get events --field-selector reason=Evicted
Debug container kubectl debug <pod> -it --image=busybox
Service endpoints kubectl get endpoints <service>

Troubleshooting Decision Tree

Single PodMultiple PodsNetworkStorageCrashLoopBackOffImagePullBackOffPendingOOMKilledYesNoProblem DetectedScope?Check Pod StatusCheck Node/ClusterTest ConnectivityCheck PVC/PVState?Logs + DescribeImage + SecretsResources + TaintsMemory LimitsNode ConditionsResource Pressure?Eviction/LimitsCordon/DrainDNS TestService/EndpointsNetwork PolicyPVC StatusStorage ClassCSI LogsSingle PodMultiple PodsNetworkStorageCrashLoopBackOffImagePullBackOffPendingOOMKilledYesNoProblem DetectedScope?Check Pod StatusCheck Node/ClusterTest ConnectivityCheck PVC/PVState?Logs + DescribeImage + SecretsResources + TaintsMemory LimitsNode ConditionsResource Pressure?Eviction/LimitsCordon/DrainDNS TestService/EndpointsNetwork PolicyPVC StatusStorage ClassCSI Logs

Resource Management Cheatsheet

Task Command
Set requests/limits kubectl set resources deployment <name> --requests=cpu=100m,memory=128Mi --limits=cpu=500m,memory=512Mi
View QoS class kubectl get pod <pod> -o jsonpath='{.status.qosClass}'
Top pods by memory kubectl top pods --sort-by=memory
Top pods by CPU kubectl top pods --sort-by=cpu
Create HPA kubectl autoscale deployment <name> --cpu-percent=80 --min=2 --max=10
View LimitRange kubectl describe limitrange
View ResourceQuota kubectl describe resourcequota
Find evicted pods kubectl get pods -A --field-selector=status.phase=Failed | grep Evicted
Delete evicted pods kubectl delete pods -A --field-selector=status.phase=Failed

Node Management Cheatsheet

Task Command
Cordon node kubectl cordon <node>
Uncordon node kubectl uncordon <node>
Drain node kubectl drain <node> --ignore-daemonsets --delete-emptydir-data
Add taint kubectl taint nodes <node> key=value:NoSchedule
Remove taint kubectl taint nodes <node> key=value:NoSchedule-
View node conditions kubectl describe node <node> | grep Conditions: -A 10
Top nodes kubectl top nodes --sort-by=memory
Node events kubectl get events --field-selector involvedObject.kind=Node

Common Issues and Solutions

Issue: Pods Stuck in ImagePullBackOff

Symptoms:

  • Pod status shows ImagePullBackOff or ErrImagePull
  • Events show "Failed to pull image"

Diagnosis:

kubectl describe pod <pod-name> | grep -A 5 "Events:"
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].image}'

Solutions:

  1. Verify image name and tag are correct
  2. Check imagePullSecrets are configured for private registries
  3. Validate registry credentials
  4. Test image pull on node manually
  5. Check network connectivity to registry

Issue: CrashLoopBackOff

Symptoms:

  • Pod repeatedly restarts
  • Status shows CrashLoopBackOff
  • Restart count increases

Diagnosis:

kubectl logs <pod-name> --previous
kubectl describe pod <pod-name>

Solutions:

  1. Check application logs for errors
  2. Verify environment variables and secrets
  3. Review liveness probe configuration
  4. Check file system permissions
  5. Validate application dependencies

Issue: Service Not Reachable

Symptoms:

  • Cannot connect to service
  • Timeout errors from clients
  • Service returns no response

Diagnosis:

kubectl get svc <service-name>
kubectl get endpoints <service-name>
kubectl describe svc <service-name>

Solutions:

  1. Verify service selector matches pod labels
  2. Check endpoints exist (pods are ready)
  3. Test connectivity from within cluster
  4. Verify port configuration
  5. Check NetworkPolicy rules

Issue: PVC Stuck in Pending

Symptoms:

  • PVC remains in Pending state
  • Pod cannot start due to volume mount failure

Diagnosis:

kubectl describe pvc <pvc-name>
kubectl get pv
kubectl get storageclass

Solutions:

  1. Create matching PersistentVolume
  2. Verify StorageClass exists and is correct
  3. Check storage provisioner is running
  4. Validate access modes match
  5. Ensure sufficient capacity available

Issue: Node NotReady

Symptoms:

  • Node status shows NotReady
  • Pods not scheduling on node
  • Node conditions show pressure

Diagnosis:

kubectl describe node <node-name>
kubectl get node <node-name> -o jsonpath='{.status.conditions}'
journalctl -u kubelet  # on node

Solutions:

  1. Check kubelet is running on node
  2. Verify network connectivity
  3. Clear disk space if DiskPressure
  4. Free memory if MemoryPressure
  5. Restart kubelet service
  6. Check CNI plugin status

Issue: DNS Resolution Failures

Symptoms:

  • Services not resolving by name
  • nslookup fails from pods
  • Application cannot find dependencies

Diagnosis:

kubectl run dnstest --image=busybox:1.28 --rm -it -- nslookup kubernetes.default
kubectl logs -n kube-system -l k8s-app=kube-dns
kubectl get svc -n kube-system kube-dns

Solutions:

  1. Verify CoreDNS pods are running
  2. Check CoreDNS ConfigMap
  3. Validate DNS service endpoint
  4. Test upstream DNS resolution
  5. Scale CoreDNS if needed

Issue: High Resource Consumption

Symptoms:

  • Pods being evicted
  • Node showing resource pressure
  • Slow cluster performance

Diagnosis:

kubectl top nodes
kubectl top pods --all-namespaces --sort-by=memory
kubectl describe node <node-name> | grep -A 20 "Allocated resources:"

Solutions:

  1. Set appropriate resource limits
  2. Implement HPA for auto-scaling
  3. Add more cluster capacity
  4. Identify and fix resource leaks
  5. Implement ResourceQuotas
  6. Use VPA for right-sizing

This comprehensive troubleshooting playbook provides systematic approaches to diagnose and resolve the most common Kubernetes issues. Regular practice with these diagnostic commands builds troubleshooting proficiency and reduces mean time to resolution (MTTR) in production environments.