Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Linux Performance Analysis

Comprehensive guide to analysing system performance, identifying bottlenecks, and optimising resource utilisation across CPU, memory, I/O, and network subsystems.

Linux Performance Analysis

Comprehensive guide to analysing system performance, identifying bottlenecks, and optimising resource utilisation across CPU, memory, I/O, and network subsystems.

Overview

Linux performance analysis involves systematic investigation of system resource utilisation to identify bottlenecks, diagnose performance degradation, and optimise workload efficiency. Modern Linux provides rich instrumentation through the kernel, including perf events, tracepoints, eBPF, and traditional tools like sar, iostat, and vmstat. Understanding these tools and metrics enables proactive monitoring, capacity planning, and rapid troubleshooting of production issues.

High CPUHigh MemorySlow I/ONetwork LatencyPerformance IssueSymptom AnalysisCPU AnalysisMemory AnalysisDisk/IO AnalysisNetwork Analysistop/htopperfpidstatfree/vmstatslabtopsmemiostatiotopblktracess/netstatiftoptcpdumpRoot CauseOptimisationHigh CPUHigh MemorySlow I/ONetwork LatencyPerformance IssueSymptom AnalysisCPU AnalysisMemory AnalysisDisk/IO AnalysisNetwork Analysistop/htopperfpidstatfree/vmstatslabtopsmemiostatiotopblktracess/netstatiftoptcpdumpRoot CauseOptimisation

CPU Performance Analysis

CPU analysis focuses on understanding utilisation patterns, identifying compute-intensive processes, and detecting scheduling inefficiencies.

Key Metrics

  • User time (us): CPU time spent running user-space processes
  • System time (sy): CPU time spent in kernel space (syscalls, kernel threads)
  • Idle (id): Percentage of time CPU is idle
  • I/O wait (wa): Time waiting for I/O completion (high values indicate I/O bottleneck)
  • Hardware interrupts (hi): Time processing hardware interrupts
  • Software interrupts (si): Time processing software interrupts (network packets, tasklets)
  • Steal time (st): Time stolen by hypervisor (virtualised environments)
  • Load average: Average number of runnable or running processes over 1, 5, 15 minutes

Real-Time Monitoring

# top - Real-time process monitor
top                                 # Interactive view
top -u username                     # Filter by user
top -p 1234,5678                    # Monitor specific PIDs
top -H                              # Show threads instead of processes
top -b -n 1                         # Batch mode, one iteration

# Key commands in top:
# 1 - Show individual CPU cores
# P - Sort by CPU usage
# M - Sort by memory usage
# k - Kill process
# r - Renice process
# f - Configure fields

# htop - Enhanced interactive viewer
htop                                # Colour-coded interactive view
htop -u username                    # Filter by user
htop -p 1234                        # Monitor specific PID
htop -d 5                           # Update delay 0.5 seconds

# mpstat - Multi-processor statistics
mpstat                              # Overall CPU stats
mpstat -P ALL                       # Per-CPU statistics
mpstat -P ALL 1 5                   # Update every 1 sec, 5 iterations
mpstat -P 0,1,2 1                   # Monitor specific CPUs

# Example output:
# CPU    %usr   %nice    %sys %iowait    %irq   %soft  %steal  %guest  %gnice   %idle
# all   12.50    0.00    3.20    0.10    0.00    0.30    0.00    0.00    0.00   83.90
#   0   25.00    0.00    5.00    0.00    0.00    0.50    0.00    0.00    0.00   69.50
#   1    8.00    0.00    2.00    0.20    0.00    0.20    0.00    0.00    0.00   89.60

Process-Level CPU Analysis

# pidstat - Per-process statistics
pidstat                             # All processes, CPU usage
pidstat 1 5                         # Update every 1 sec, 5 times
pidstat -p 1234                     # Specific process
pidstat -u                          # CPU statistics (default)
pidstat -C nginx                    # Filter by command name
pidstat -t                          # Show threads
pidstat -T ALL                      # Show all tasks

# Example output:
# Linux 5.15.0 (hostname)    12/07/2025    _x86_64_    (8 CPU)
#
# 10:15:23 AM   UID       PID    %usr %system  %guest   %wait    %CPU   CPU  Command
# 10:15:24 AM  1000      1234   25.00    5.00    0.00    0.00   30.00     0  nginx
# 10:15:24 AM  1000      5678   12.00    2.00    0.00    1.00   14.00     2  python3

# ps - Process snapshot
ps aux --sort=-%cpu | head -20      # Top CPU consumers
ps -eo pid,ppid,cmd,%cpu,%mem --sort=-%cpu | head -20
ps -eo pid,tid,class,rtprio,ni,pri,psr,pcpu,stat,wchan:14,comm
ps -p 1234 -L                       # Show threads for PID

# Check CPU affinity
taskset -cp 1234                    # Show CPU affinity for PID

# Set CPU affinity
taskset -cp 0,1,2,3 1234            # Pin process to CPUs 0-3

Historical CPU Analysis

# sar - System Activity Reporter
sar -u                              # CPU utilisation (today)
sar -u 1 5                          # Live: every 1 sec, 5 times
sar -u -f /var/log/sa/sa07          # Historical data from specific day
sar -u -s 10:00:00 -e 12:00:00      # Time range
sar -P ALL                          # Per-CPU statistics

# Example output:
# 10:00:01 AM     CPU     %user     %nice   %system   %iowait    %steal     %idle
# 10:10:01 AM     all     15.23      0.00      3.45      0.12      0.00     81.20
# Average:        all     14.87      0.00      3.52      0.15      0.00     81.46

# uptime - Load averages
uptime                              # Current load averages
# Output: 10:15:23 up 45 days, 2:34, 3 users, load average: 2.45, 1.89, 1.56
# Load > CPU count suggests queuing; <1 per CPU is healthy. Note: on Linux,
# load also counts uninterruptible (D-state) I/O waiters, so heavy disk/NFS
# can inflate it without the CPUs being the bottleneck — corroborate with %wa.

# Check run queue
sar -q 1 5                          # Run queue and load average
vmstat 1 5                          # 'r' column shows runnable processes

CPU Context Switches and Interrupts

# vmstat - Virtual memory statistics
vmstat 1 5                          # 1 second intervals, 5 samples
vmstat -w 1                         # Wide output format

# Key columns:
# r  - Runnable processes (run queue)
# b  - Blocked processes (uninterruptible sleep)
# cs - Context switches per second
# in - Interrupts per second

# Example output:
# procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
#  r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
#  2  0      0 524288  76800 819200    0    0     5    10  250  500 12  3 85  0  0

# A sustained context-switch rate well above your baseline (judge per-core,
# not as an absolute — a busy I/O box can run 100k+/sec quite happily) may indicate:
# - Too many threads
# - Lock contention
# - Inefficient scheduling

# pidstat for context switches
pidstat -w 1 5                      # Context switch statistics
pidstat -w -p 1234 1                # Specific process

# Example output:
# 10:15:23 AM   UID       PID   cswch/s nvcswch/s  Command
# 10:15:24 AM  1000      1234     50.00    100.00  nginx
# cswch   - Voluntary context switches (I/O wait)
# nvcswch - Non-voluntary (preempted by scheduler)

CPU Frequency and Power States

# Check CPU frequency scaling
cat /proc/cpuinfo | grep MHz
lscpu | grep MHz

# cpupower - CPU power management
cpupower frequency-info             # Current frequency settings
cpupower frequency-set -g performance   # Set governor to performance
cpupower idle-info                  # C-state information

# turbostat - Advanced CPU statistics
turbostat --interval 1              # Real-time frequency, C-states, temperature

Memory Performance Analysis

Memory analysis involves understanding allocation patterns, detecting leaks, and identifying swap pressure and page cache efficiency.

Linux Memory ArchitectureYesNoPhysical RAMKernel SpaceUser SpaceSlab CachesPage CacheKernel BuffersProcess MemoryShared LibrariesFile MappingsVirtual MemoryAnonymous PagesFile-backed PagesMemory Pressure?SwapResident MemoryLinux Memory ArchitectureYesNoPhysical RAMKernel SpaceUser SpaceSlab CachesPage CacheKernel BuffersProcess MemoryShared LibrariesFile MappingsVirtual MemoryAnonymous PagesFile-backed PagesMemory Pressure?SwapResident Memory

Key Metrics

  • Total memory: Physical RAM available
  • Used: Memory in use (total - free - buffers - cache)
  • Free: Unallocated memory (low values normal due to caching)
  • Available: Memory available without swapping
  • Buffers: Filesystem metadata cache
  • Cached: Page cache for file data
  • Swap: Disk-based overflow memory
  • RSS (Resident Set Size): Physical memory used by process
  • VSZ (Virtual Size): Virtual memory allocated (including swapped)
  • Dirty pages: Modified pages pending write to disk

System-Wide Memory Analysis

# free - Memory overview
free                                # Basic memory stats
free -h                             # Human-readable (GB, MB)
free -m                             # In megabytes
free -g                             # In gigabytes
free -s 1                           # Update every 1 second
free -w                             # Wide format (buffers/cache separate)

# Example output:
#               total        used        free      shared  buff/cache   available
# Mem:           15Gi       8.0Gi       2.0Gi       500Mi       5.0Gi       6.5Gi
# Swap:         2.0Gi       100Mi       1.9Gi

# /proc/meminfo - Detailed memory information
cat /proc/meminfo                   # Comprehensive memory details
grep -E 'MemTotal|MemFree|MemAvailable|Buffers|Cached|SwapTotal|SwapFree|Dirty|Writeback' /proc/meminfo

# vmstat - Memory statistics
vmstat 1 5                          # Memory + swap activity

# Key columns:
# swpd - Swap used
# free - Free memory
# buff - Buffer cache
# cache - Page cache
# si - Swap in (from disk)
# so - Swap out (to disk)
# High si/so indicates memory pressure

Process Memory Analysis

# ps - Process memory usage
ps aux --sort=-%mem | head -20      # Top memory consumers
ps -eo pid,ppid,cmd,rss,vsz,%mem --sort=-rss | head -20

# Column explanations:
# RSS  - Resident Set Size (physical memory)
# VSZ  - Virtual Size (total virtual memory)
# %MEM - Percentage of physical memory

# pidstat - Per-process memory stats
pidstat -r                          # Memory statistics
pidstat -r 1 5                      # Every 1 sec, 5 times
pidstat -r -p 1234                  # Specific process

# Example output:
# 10:15:23 AM   UID       PID  minflt/s  majflt/s     VSZ     RSS   %MEM  Command
# 10:15:24 AM  1000      1234    150.00      0.00  524288  262144   1.60  nginx
# minflt - Minor page faults (page in cache)
# majflt - Major page faults (read from disk)

# smem - Memory reporting tool
smem                                # Per-process memory usage
smem -t                             # Show totals
smem -u                             # Per-user summary
smem -m                             # System-wide summary
smem -s uss                         # Sort by unique set size
smem -p                             # Show percentages

# pmap - Process memory map
pmap 1234                           # Memory map for PID
pmap -x 1234                        # Extended format
pmap -X 1234                        # More details
pmap -XX 1234                       # Everything

# /proc/<pid>/status - Detailed process memory
grep -E 'VmSize|VmRSS|VmData|VmStk|VmExe|VmLib' /proc/1234/status

# /proc/<pid>/smaps - Detailed memory mappings
cat /proc/1234/smaps                # Full memory mapping details
cat /proc/1234/smaps_rollup         # Summary only

Memory Pressure and OOM

# Check for OOM kills
dmesg | grep -i oom                 # OOM killer messages
journalctl -k | grep -i oom         # Systemd journal

# Example OOM message (modern kernels):
# Out of memory: Killed process 1234 (process_name) total-vm:2097152kB, \
#   anon-rss:1048576kB, file-rss:0kB, shmem-rss:0kB, UID:1000 \
#   pgtables:4096kB oom_score_adj:0

# Monitor OOM score
cat /proc/1234/oom_score             # Current OOM score (higher = more likely to kill)
cat /proc/1234/oom_score_adj         # OOM adjustment (-1000 to 1000)

# Adjust OOM score to protect critical processes
echo -1000 > /proc/1234/oom_score_adj   # Never kill
echo -500 > /proc/1234/oom_score_adj    # Lower priority for OOM

# vm.swappiness - Control swap aggressiveness
sysctl vm.swappiness                # Current value (0-100)
sysctl -w vm.swappiness=10          # Lower value = less swap (prefer RAM)
# 0   - Avoid swap entirely until OOM is imminent (can trigger OOM kills
#       where you expected "just a little swap" — a common footgun)
# 1   - Safest true minimum: still permits swapping under real pressure
# 10  - Minimal swap (recommended for servers)
# 60  - Default
# 100 - Aggressive swap

# Memory pressure stall information (PSI)
cat /proc/pressure/memory           # Memory pressure metrics
# some avg10=0.00 avg60=0.05 avg300=0.10 total=1000000
# full avg10=0.00 avg60=0.00 avg300=0.00 total=500000

Slab Cache Analysis

# slabtop - Kernel slab cache statistics
slabtop                             # Real-time slab cache view
slabtop -o                          # Display once and exit (default sort: active objects)
slabtop -s c                        # Sort by cache size

# /proc/slabinfo - Detailed slab information
cat /proc/slabinfo                  # Raw slab cache data
cat /proc/slabinfo | head -20

# Example output shows kernel object caches:
# dentry (directory entries), inode_cache, buffer_head, etc.

# Check for slab memory leaks
watch -n 1 'cat /proc/slabinfo | grep dentry'

Page Cache Analysis

# Check page cache usage
free -h                             # 'buff/cache' column

# Drop page caches (requires root)
sync                                # Flush pending writes
echo 1 > /proc/sys/vm/drop_caches   # Drop page cache only
echo 2 > /proc/sys/vm/drop_caches   # Drop dentries and inodes
echo 3 > /proc/sys/vm/drop_caches   # Drop everything
# Note: Only for testing; kernel manages caches efficiently

# Check dirty page writeback
cat /proc/sys/vm/dirty_ratio        # % of RAM before sync write
cat /proc/sys/vm/dirty_background_ratio  # % of RAM before async write
cat /proc/meminfo | grep Dirty      # Current dirty pages

I/O and Disk Performance Analysis

I/O analysis focuses on disk throughput, latency, queue depths, and identifying bottlenecks in storage subsystems.

Key Metrics

  • IOPS: Input/output operations per second
  • Throughput: Data transfer rate (MB/s, GB/s)
  • Latency: Time to complete I/O operation (ms)
  • Queue depth: Number of pending I/O requests
  • Utilisation: Percentage of time device is busy
  • Service time: Average time to service requests
  • Wait time: Time requests spend in queue
  • %util: Device utilisation (100% indicates saturation)

System-Wide I/O Statistics

# iostat - I/O statistics
iostat                              # Basic CPU and device stats
iostat -x                           # Extended statistics
iostat -x 1 5                       # Every 1 sec, 5 times
iostat -x -d                        # Devices only (no CPU)
iostat -x sda sdb                   # Specific devices
iostat -x -m                        # In MB/s instead of KB/s
iostat -x -p ALL                    # Include partitions

# Example extended output (sysstat 12.x):
# Device   r/s   rkB/s rrqm/s %rrqm r_await rareq-sz   w/s  wkB/s w_await wareq-sz aqu-sz %util
# nvme0n1 50.0  2000.0   1.0   2.0    0.30    40.00  30.0 1500.0    0.50    50.00   0.03  12.0
#
# Key columns (sysstat 12+ renamed these; avgrq-sz/avgqu-sz/svctm are gone):
# r/s, w/s          - Reads/writes per second
# rkB/s, wkB/s      - Read/write throughput
# rrqm/s, wrqm/s    - Read/write requests merged per second
# r_await, w_await  - Average read/write wait time (ms) - queue + service time
# rareq-sz, wareq-sz- Average read/write request size (kB), replaces avgrq-sz
# aqu-sz            - Average queue length, replaces avgqu-sz
# %util             - Utilisation % (100% = saturation)
# (svctm and a single combined await were removed; there is no service-time column)

# sar - Historical I/O statistics
sar -b                              # Overall I/O transfer stats
sar -b 1 5                          # Live stats
sar -d                              # Block device statistics
sar -d -p                           # Pretty device names

# Example output:
# 10:15:23 AM       tps      rtps      wtps   bread/s   bwrtn/s
# 10:15:24 AM    100.00     50.00     50.00   2048.00   1024.00
# tps  - Transactions per second
# rtps - Read transactions per second
# wtps - Write transactions per second

Process I/O Analysis

# iotop - I/O by process (requires root)
iotop                               # Interactive I/O monitor
iotop -o                            # Only show processes doing I/O
iotop -a                            # Accumulated I/O
iotop -P                            # Show processes (not threads)
iotop -p 1234                       # Monitor specific PID

# pidstat - Per-process I/O statistics
pidstat -d                          # Disk I/O statistics
pidstat -d 1 5                      # Every 1 sec, 5 times
pidstat -d -p 1234                  # Specific process

# Example output:
# 10:15:23 AM   UID       PID   kB_rd/s   kB_wr/s kB_ccwr/s iodelay  Command
# 10:15:24 AM  1000      1234    500.00   1000.00      0.00      10  mysqld
# kB_rd/s  - Kilobytes read per second
# kB_wr/s  - Kilobytes written per second
# kB_ccwr/s - Cancelled writes
# iodelay  - Block I/O delay (clock ticks)

# /proc/<pid>/io - Process I/O counters
cat /proc/1234/io
# rchar, wchar   - Total bytes read/written (includes cache)
# syscr, syscw   - Read/write syscalls
# read_bytes, write_bytes - Actual disk I/O
# cancelled_write_bytes - Truncated dirty cache

# Find processes with most I/O
# (stock ps has no block-I/O sort key — use pidstat -d or iotop above instead)
pidstat -d 1 1 | sort -k6 -rn | head  # Sort by kB_wr/s

Open File Analysis with lsof

lsof (list open files) maps file descriptors to processes — invaluable for finding what holds a busy mount, leaks descriptors, or pins disk space via deleted-but-open files.

# What has a file or directory open
lsof /var/log/app.log               # Processes with this file open
lsof +D /mnt/data                   # Everything open under a directory (recursive, slow)
lsof +d /mnt/data                   # One level only (faster than +D)

# Why a filesystem won't unmount ("target is busy")
lsof /mnt/data                      # Find the offenders before umount
fuser -vm /mnt/data                 # Lighter alternative; -k would kill them

# Per-process open files
lsof -p 1234                        # All FDs for a PID
lsof -p 1234 | wc -l                # Crude FD count (watch for descriptor leaks)
ls /proc/1234/fd | wc -l            # Exact FD count from /proc
lsof -u www-data                    # All files opened by a user

# Network sockets (lsof doubles as a socket lister)
lsof -i                             # All network connections
lsof -i :443                        # What is listening on / using port 443
lsof -iTCP -sTCP:LISTEN             # Listening TCP sockets only
lsof -i -n -P                       # Skip DNS/port-name lookups (faster, raw numbers)

# Disk space held by deleted-but-open files (df full, du doesn't agree)
lsof -nP +L1                        # Files with link count 0 still held open
lsof | grep '(deleted)'             # Classic form; space frees when the FD closes

Tip: if df shows a disk full but du can't find the space, a process is still holding a deleted file open. Find it with lsof -nP +L1, then restart or signal that process to release the descriptor.

Block Device Analysis

# lsblk - List block devices
lsblk                               # Tree view of devices
lsblk -f                            # Include filesystem info
lsblk -o NAME,SIZE,TYPE,MOUNTPOINT,FSTYPE,MODEL

# Check I/O scheduler (brackets mark the active one)
cat /sys/block/nvme0n1/queue/scheduler
# Typical output on NVMe: [none] mq-deadline
# Common schedulers:
# none        - No scheduler; default for NVMe (low-latency, deep hardware queues)
# mq-deadline - Default for SATA/SAS SSDs and HDDs
# bfq         - Budget Fair Queueing (desktop interactivity; may need kernel module)
# kyber       - Low-latency scheduler (may need kernel module)

# Change I/O scheduler
echo mq-deadline > /sys/block/sda/queue/scheduler

# Check queue depth
cat /sys/block/sda/queue/nr_requests    # Current queue size
cat /sys/block/sda/queue/scheduler      # Active scheduler

# Device statistics
cat /proc/diskstats                 # Raw device statistics
cat /sys/block/sda/stat             # Per-device stats

Advanced I/O Tracing

# blktrace - Block layer tracing
blktrace -d /dev/sda -o - | blkparse -i -    # Real-time trace
blktrace -d /dev/sda                         # Capture to files
blkparse sda                                 # Parse captured trace

# iowatcher - Visualise blktrace data
blktrace -d /dev/sda -o sda
iowatcher -t sda.blktrace.* -o trace.svg     # Generate graph

# ftrace - Function tracing for I/O
cd /sys/kernel/debug/tracing
echo blk > current_tracer
echo 1 > tracing_on
cat trace
echo 0 > tracing_on

Filesystem Performance

# df - Disk space and inode usage
df -h                               # Human-readable disk space
df -i                               # Inode usage
df -T                               # Include filesystem type

# Check mount options affecting performance
mount | grep -E 'sda|nvme'
# Look for: noatime, nodiratime, relatime, barrier, journal mode

# Filesystem statistics
# For ext4:
tune2fs -l /dev/sda1 | grep -E 'Block count|Free blocks|Inode count'

# For XFS:
xfs_info /mount/point

# Check for filesystem errors
dmesg | grep -E 'EXT4-fs|XFS'
journalctl -k | grep -E 'filesystem|I/O error'

Network Performance Analysis

Network analysis examines throughput, latency, packet loss, connection states, and protocol-level performance.

Key Metrics

  • Throughput: Data transfer rate (Mbps, Gbps)
  • Packet rate: Packets per second
  • Latency: Round-trip time (RTT)
  • Errors: Transmit/receive errors, drops
  • Retransmissions: TCP segments retransmitted
  • Connection states: Established, time-wait, close-wait
  • Buffer usage: Socket send/receive buffer occupancy

Network Interface Statistics

# ip - Network interface statistics
ip -s link                          # Interface statistics
ip -s -s link                       # More detailed statistics
ip -s link show eth0                # Specific interface

# Example output (iproute2; RX shows 'missed' not 'overrun' on modern kernels):
# 2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500
#     RX: bytes  packets  errors  dropped missed  mcast
#     1048576000 5000000  0       5       0       100
#     TX: bytes  packets  errors  dropped carrier collsns
#     524288000  3000000  0       0       0       0
# ip -s -s link adds detailed RX/TX error breakdowns (length, crc, frame, fifo, overrun)

# ifconfig - Legacy interface stats
ifconfig eth0                       # Interface statistics

# netstat - Network statistics (deprecated, use ss)
netstat -i                          # Interface statistics
netstat -s                          # Protocol statistics

# sar - Network statistics
sar -n DEV                          # Network device statistics
sar -n DEV 1 5                      # Live stats
sar -n EDEV                         # Network errors
sar -n TCP                          # TCP statistics
sar -n SOCK                         # Socket statistics

# Example DEV output:
# 10:15:23 AM     IFACE   rxpck/s   txpck/s    rxkB/s    txkB/s
# 10:15:24 AM      eth0   5000.00   3000.00   2500.00   1500.00

Connection Analysis

# ss - Socket statistics (modern netstat replacement)
ss -s                               # Summary statistics
ss -t                               # TCP connections
ss -u                               # UDP connections
ss -l                               # Listening sockets
ss -a                               # All sockets
ss -p                               # Show process using socket
ss -n                               # Numeric (no DNS resolution)
ss -e                               # Extended info
ss -m                               # Memory usage

# Common combinations:
ss -tulpn                           # TCP/UDP listening, process, numeric
ss -tan                             # All TCP, numeric
ss -tan state established           # Established TCP connections

# Connection state filtering:
ss state established                # Established connections
ss state time-wait                  # Time-wait connections
ss state close-wait                 # Close-wait (app not closing properly)

# Advanced filtering:
ss dst 192.168.1.100                # Connections to specific host
ss sport = :80                      # Connections from port 80
ss dport = :443                     # Connections to port 443
ss -o state established '( dport = :22 or sport = :22 )'    # SSH connections

# Connection statistics by state:
ss -tan | awk '{print $1}' | sort | uniq -c | sort -rn

# Monitor connection rate:
watch -n 1 'ss -tan | wc -l'

Network Throughput and Bandwidth

# iftop - Real-time bandwidth usage by connection (requires root)
iftop                               # Default interface
iftop -i eth0                       # Specific interface
iftop -n                            # No DNS resolution
iftop -P                            # Show ports
iftop -B                            # Display in bytes

# nload - Network load visualisation
nload                               # All interfaces
nload eth0                          # Specific interface
nload -u M                          # Units in Mbps

# bmon - Bandwidth monitor
bmon                                # Interactive bandwidth monitor
bmon -p eth0                        # Specific interface

# nethogs - Network usage by process (requires root)
nethogs                             # Default interface
nethogs eth0                        # Specific interface

# iperf3 - Network performance testing
# On server:
iperf3 -s                           # Start server
# On client:
iperf3 -c server_ip                 # TCP test
iperf3 -c server_ip -u -b 1G        # UDP test at 1 Gbps
iperf3 -c server_ip -P 4            # 4 parallel streams
iperf3 -c server_ip -t 60           # 60 second test

Protocol-Level Analysis

# TCP statistics
ss -ti                              # TCP info for established connections
ss -ti state established            # With connection state filter

# Example output shows:
# - RTT (Round-trip time)
# - Congestion window (cwnd)
# - Send and receive buffers
# - Retransmissions

# netstat protocol stats (legacy)
netstat -s                          # All protocol statistics
netstat -st                         # TCP statistics only
netstat -su                         # UDP statistics only

# Look for:
# - TCP retransmissions
# - Segment receive errors
# - UDP packet receive errors

# tcplife - Trace TCP session lifespans (BCC/eBPF)
# On Debian/Ubuntu the bpfcc-tools package suffixes these as -bpfcc
tcplife-bpfcc                       # Show TCP connections
tcplife-bpfcc -p 1234               # Filter by PID

# tcptop - Top TCP throughput by connection
tcptop-bpfcc                        # Real-time TCP throughput
tcptop-bpfcc -p 1234                # Filter by PID

Packet Capture and Analysis

# tcpdump - Packet capture
tcpdump -i eth0                     # Capture on interface
tcpdump -i any                      # Capture on all interfaces
tcpdump -i eth0 -n                  # No DNS resolution
tcpdump -i eth0 -c 100              # Capture 100 packets
tcpdump -i eth0 -w capture.pcap     # Write to file
tcpdump -r capture.pcap             # Read from file

# Filtering:
tcpdump -i eth0 port 80             # HTTP traffic
tcpdump -i eth0 host 192.168.1.100  # Specific host
tcpdump -i eth0 tcp                 # TCP only
tcpdump -i eth0 'tcp[tcpflags] & tcp-syn != 0'    # SYN packets

# Check for packet loss:
tcpdump -i eth0 -vv -c 1000 2>&1 | grep 'packets dropped by kernel'

Network Latency Analysis

# ping - ICMP latency test
ping -c 10 host                     # 10 packets
ping -i 0.2 -c 50 host              # 0.2s interval, 50 packets
ping -s 1500 host                   # Jumbo frame test

# mtr - Network diagnostic tool (traceroute + ping)
mtr host                            # Interactive
mtr --report host                   # Report mode
mtr --report -c 100 host            # 100 packets

# ss with RTT information
ss -ti | grep rtt                   # Check TCP RTT
ss -ti dst 192.168.1.100 | grep rtt # RTT to specific host

Profiling with perf and Flamegraphs

The perf tool provides low-overhead CPU profiling using hardware performance counters and software events.

Applicationperf recordperf.dataperf reportperf scriptflamegraph.plFlamegraph SVGText ReportVisual AnalysisApplicationperf recordperf.dataperf reportperf scriptflamegraph.plFlamegraph SVGText ReportVisual Analysis

Key Concepts

  • Hardware counters: CPU cycle counts, cache misses, branch mispredictions
  • Software events: Page faults, context switches, CPU migrations
  • Tracepoints: Kernel static instrumentation points
  • Dynamic probes: Kprobes, uprobes for kernel and userspace
  • Sampling: Periodic snapshots of call stacks
  • Flamegraphs: Visual representation of CPU time by stack trace

Basic Profiling

# perf top - Real-time profiling
perf top                            # System-wide top functions
perf top -p 1234                    # Specific process
perf top -g                         # Show call graphs
perf top -e cycles                  # Profile CPU cycles
perf top -e cache-misses            # Profile cache misses

# perf record - Record profile data
perf record -a                      # System-wide
perf record -p 1234                 # Specific process
perf record -g                      # Record call graphs
perf record -F 99                   # Sample at 99 Hz
perf record -a -g sleep 10          # Record for 10 seconds
perf record -e cpu-clock -g -p 1234 # CPU clock events for PID

# perf report - Analyse recorded data
perf report                         # Interactive report
perf report --stdio                 # Text output
perf report -g                      # Show call graphs
perf report --sort comm,dso         # Sort by command and DSO

# Navigation in perf report:
# Enter - Drill down into function
# ESC   - Go back
# /     - Search
# a     - Annotate (show assembly)

Advanced Profiling

# Profile specific events
perf list                           # List all available events
perf list cache                     # List cache events
perf list 'block:*'                 # List block layer tracepoints

# Common hardware events:
perf record -e cycles -a -g sleep 10              # CPU cycles
perf record -e instructions -a -g sleep 10        # Instructions
perf record -e cache-references -a -g sleep 10    # Cache refs
perf record -e cache-misses -a -g sleep 10        # Cache misses
perf record -e branch-misses -a -g sleep 10       # Branch mispredictions
perf record -e LLC-loads -a -g sleep 10           # Last-level cache loads
perf record -e LLC-load-misses -a -g sleep 10     # LLC misses

# Multiple events:
perf record -e cycles,instructions,cache-misses -a -g sleep 10

# perf stat - Event statistics
perf stat command                   # Basic stats for command
perf stat -a sleep 10               # System-wide for 10 seconds
perf stat -p 1234 sleep 10          # Specific process for 10 seconds
perf stat -e cycles,instructions,cache-misses command

# Example output:
#  Performance counter stats for 'command':
#
#     10,234,567,890      cycles
#      8,123,456,789      instructions              #    0.79  insn per cycle
#        123,456,789      cache-misses              #   12.34 % of all cache refs

Call Graph Analysis

# Record with different call graph methods
perf record --call-graph fp command       # Frame pointer (fast, least accurate)
perf record --call-graph dwarf command    # DWARF (accurate, large data)
perf record --call-graph lbr command      # Last Branch Record (Intel)
# Note: bare -g defaults to fp; the method must follow --call-graph, not -g
# (i.e. "perf record -g fp command" treats fp as the workload and fails)

# Analyse call graphs
perf report -g 'graph,0.5,caller'   # Show call graph with 0.5% threshold
perf report -g 'fractal'            # Fractal call graph view
perf report --no-children           # Hide children overhead

# Generate call graph data for external tools
perf script > out.perf              # Export for processing

Generating Flamegraphs

# Install flamegraph tools
git clone https://github.com/brendangregg/FlameGraph
cd FlameGraph

# Record profile with call stacks
perf record -F 99 -a -g -- sleep 60

# Generate flamegraph
perf script | ./stackcollapse-perf.pl | ./flamegraph.pl > flamegraph.svg

# For specific process:
perf record -F 99 -p 1234 -g -- sleep 60
perf script | ./stackcollapse-perf.pl | ./flamegraph.pl > flamegraph.svg

# Off-CPU flamegraph (blocking):
# Requires BCC tools (offcputime-bpfcc on Debian/Ubuntu; offcputime elsewhere)
offcputime-bpfcc -p 1234 60 -f > out.stacks
./flamegraph.pl --color=io --title="Off-CPU Time" out.stacks > offcpu.svg

# Differential flamegraphs:
perf record -F 99 -a -g -o perf1.data -- sleep 30   # Before optimisation
perf record -F 99 -a -g -o perf2.data -- sleep 30   # After optimisation
perf script -i perf1.data | ./stackcollapse-perf.pl > out1.folded
perf script -i perf2.data | ./stackcollapse-perf.pl > out2.folded
./difffolded.pl out1.folded out2.folded | ./flamegraph.pl > diff.svg

eBPF-Based Profiling

# BCC tools (requires bpfcc-tools package)
# On Debian/Ubuntu the commands carry a -bpfcc suffix (e.g. profile-bpfcc).
# On other distros (RHEL/Fedora) they are unsuffixed (profile, offcputime, ...).
# CPU profiling:
profile-bpfcc 99                    # Sample at 99 Hz
profile-bpfcc -F 199 -p 1234 10     # 199 Hz, PID 1234, 10 seconds
profile-bpfcc -adf -p 1234 10       # User+kernel stacks, folded output

# Off-CPU analysis:
offcputime-bpfcc 10                 # Sample off-CPU time for 10s
offcputime-bpfcc -p 1234 -u 10      # User stacks only

# Function latency:
funclatency-bpfcc do_sys_open       # Latency histogram for kernel function
funclatency-bpfcc -p 1234 'c:read'  # Latency for libc read() in PID

# Stack sampling:
stackcount-bpfcc -p 1234 -r '^tcp_sendmsg'    # Count TCP send stacks

cgroups and Resource Limiting

Control groups (cgroups) enable resource isolation and limiting for processes, essential for containerisation and multi-tenancy.

cgroup v2 HierarchyCPU limitMemory limitI/O limitRoot cgroup /system.sliceuser.slicecustom.sliceservice1service2user-1000.slicecontainer1container2cpu.maxmemory.maxio.maxcgroup v2 HierarchyCPU limitMemory limitI/O limitRoot cgroup /system.sliceuser.slicecustom.sliceservice1service2user-1000.slicecontainer1container2cpu.maxmemory.maxio.max

cgroup Versions

  • cgroup v1: Original implementation, separate hierarchies per controller
  • cgroup v2: Unified hierarchy, single tree for all controllers
# Check cgroup version
mount | grep cgroup
stat -fc %T /sys/fs/cgroup
# Output:
# cgroup2fs - cgroup v2
# tmpfs     - cgroup v1

# cgroup controllers
cat /sys/fs/cgroup/cgroup.controllers    # Available controllers (v2)
# cpu io memory pids rdma

CPU Limiting with cgroups

# cgroup v2 CPU control
# Create a cgroup
mkdir /sys/fs/cgroup/myapp

# Set CPU quota (1 CPU = 100000)
echo "100000 1000000" > /sys/fs/cgroup/myapp/cpu.max
# Format: $MAX $PERIOD  (default cpu.max is "max 100000" — a 100ms period)
# 100000/1000000 = 1 CPU core
# 50000/1000000  = 0.5 CPU cores
# max/1000000    = No limit

# Set CPU weight (shares)
echo 200 > /sys/fs/cgroup/myapp/cpu.weight
# Default: 100, Range: 1-10000
# Higher weight = more CPU time when contention exists

# Add process to cgroup
echo 1234 > /sys/fs/cgroup/myapp/cgroup.procs

# View CPU statistics
cat /sys/fs/cgroup/myapp/cpu.stat
# usage_usec - Total CPU time in microseconds
# user_usec  - User time
# system_usec - System time
# nr_periods - Number of enforcement periods
# nr_throttled - Number of times throttled
# throttled_usec - Total throttled time

# Monitor throttling
watch -n 1 cat /sys/fs/cgroup/myapp/cpu.stat

Memory Limiting with cgroups

# cgroup v2 memory control
# Set memory limit
echo "1G" > /sys/fs/cgroup/myapp/memory.max
echo "512M" > /sys/fs/cgroup/myapp/memory.high    # Soft limit (throttle)
echo "2G" > /sys/fs/cgroup/myapp/memory.swap.max  # Swap limit

# View memory usage
cat /sys/fs/cgroup/myapp/memory.current           # Current usage
cat /sys/fs/cgroup/myapp/memory.stat              # Detailed stats

# Key memory.stat fields:
# anon      - Anonymous memory (heap, stack)
# file      - Page cache
# kernel_stack - Kernel stack memory
# slab      - Kernel slab memory
# sock      - Socket buffer memory
# shmem     - Shared memory

# Memory events (OOM, limits hit)
cat /sys/fs/cgroup/myapp/memory.events
# low       - Hit memory.low threshold
# high      - Hit memory.high threshold
# max       - Hit memory.max threshold
# oom       - OOM killer invoked
# oom_kill  - Processes killed

# Set memory protection
echo "256M" > /sys/fs/cgroup/myapp/memory.min     # Hard protection
echo "512M" > /sys/fs/cgroup/myapp/memory.low     # Best-effort protection

I/O Limiting with cgroups

# cgroup v2 I/O control
# Get device major:minor numbers (the 8:0 below is SATA sda; an
# NVMe device is major 259, e.g. 259:0 — always check your own)
lsblk -o NAME,MAJ:MIN
# sda       8:0       (SATA/SAS)
# nvme0n1 259:0       (NVMe)

# Set I/O limits
# Format: $MAJ:$MIN rbps=$READ_BPS wbps=$WRITE_BPS riops=$READ_IOPS wiops=$WRITE_IOPS
echo "8:0 rbps=10485760 wbps=10485760" > /sys/fs/cgroup/myapp/io.max
# 10485760 bytes/s = 10 MB/s read and write

# IOPS limits:
echo "8:0 riops=1000 wiops=500" > /sys/fs/cgroup/myapp/io.max

# Combined:
echo "8:0 rbps=104857600 wbps=52428800 riops=10000 wiops=5000" > /sys/fs/cgroup/myapp/io.max

# Set I/O weight (proportional)
echo "8:0 100" > /sys/fs/cgroup/myapp/io.weight
# Default: 100, Range: 1-10000

# View I/O statistics
cat /sys/fs/cgroup/myapp/io.stat
# 8:0 rbytes=1048576000 wbytes=524288000 rios=10000 wios=5000 dbytes=0 dios=0

PID Limiting

# Limit number of processes/threads
echo 100 > /sys/fs/cgroup/myapp/pids.max

# View current process count
cat /sys/fs/cgroup/myapp/pids.current

# View events
cat /sys/fs/cgroup/myapp/pids.events
# max - Number of times pids.max was hit

Systemd cgroup Integration

# Systemd automatically creates cgroups for services
# View service cgroup
systemctl status nginx
# Shows: CGroup: /system.slice/nginx.service

# Set CPU limit via systemd
systemctl set-property nginx.service CPUQuota=50%    # 0.5 CPU
systemctl set-property nginx.service CPUWeight=200   # CPU shares

# Set memory limit
systemctl set-property nginx.service MemoryMax=1G
systemctl set-property nginx.service MemoryHigh=512M

# Set I/O limits
systemctl set-property nginx.service IOReadBandwidthMax="/dev/sda 10M"
systemctl set-property nginx.service IOWriteBandwidthMax="/dev/sda 5M"

# Make changes persistent
systemctl set-property --runtime nginx.service CPUQuota=50%    # Temporary
systemctl set-property nginx.service CPUQuota=50%              # Persistent

# View cgroup properties
systemctl show nginx.service -p CPUQuota -p MemoryMax

Monitoring cgroup Usage

# systemd-cgtop - cgroup resource usage
systemd-cgtop                       # Real-time cgroup monitor
systemd-cgtop -m                    # Sort by memory
systemd-cgtop -c                    # Sort by CPU
systemd-cgtop --depth=3             # Hierarchy depth

# Show all processes in a cgroup
cat /sys/fs/cgroup/myapp/cgroup.procs

# Show threads
cat /sys/fs/cgroup/myapp/cgroup.threads

# Recursively view cgroup tree
systemd-cgls

# View specific service cgroup
systemd-cgls /system.slice/nginx.service

Kernel Tuning with sysctl

The sysctl interface allows runtime modification of kernel parameters for performance tuning.

Network Tuning

# View current network parameters
sysctl -a | grep net.

# TCP buffer sizes
sysctl net.ipv4.tcp_rmem            # TCP receive buffer min,default,max
sysctl net.ipv4.tcp_wmem            # TCP send buffer min,default,max
sysctl net.core.rmem_max            # Max receive buffer size
sysctl net.core.wmem_max            # Max send buffer size

# Recommended high-performance settings:
sysctl -w net.core.rmem_max=134217728           # 128 MB
sysctl -w net.core.wmem_max=134217728           # 128 MB
sysctl -w net.ipv4.tcp_rmem="4096 87380 67108864"    # 64 MB max
sysctl -w net.ipv4.tcp_wmem="4096 65536 67108864"    # 64 MB max

# TCP congestion control
sysctl net.ipv4.tcp_congestion_control          # Current algorithm
sysctl net.ipv4.tcp_available_congestion_control  # List loaded algorithms
sysctl -w net.ipv4.tcp_congestion_control=bbr   # Use BBR
# Options: cubic (default), bbr, reno, vegas
# reno and cubic are built in; bbr/vegas may need: modprobe tcp_bbr

# TCP connection settings
sysctl -w net.ipv4.tcp_max_syn_backlog=8192     # SYN queue size
sysctl -w net.core.somaxconn=4096               # Accept queue size
sysctl -w net.ipv4.tcp_fin_timeout=30           # FIN_WAIT2 timeout
sysctl -w net.ipv4.tcp_tw_reuse=1               # Reuse TIME_WAIT sockets (outbound); kernel default is now 2 (loopback only)
sysctl -w net.ipv4.tcp_slow_start_after_idle=0  # Disable slow start

# Connection tracking
sysctl -w net.netfilter.nf_conntrack_max=1048576    # Max tracked connections
sysctl -w net.nf_conntrack_max=1048576              # Alternative path

# Network queue lengths
sysctl -w net.core.netdev_max_backlog=5000      # Incoming packet queue
sysctl -w net.core.netdev_budget=600            # Packets per NAPI poll

Virtual Memory Tuning

# Swappiness (tendency to swap)
sysctl vm.swappiness                # View current (default: 60)
sysctl -w vm.swappiness=10          # Prefer RAM over swap

# Dirty page writeback
sysctl vm.dirty_ratio               # % RAM before sync write (default: 20)
sysctl vm.dirty_background_ratio    # % RAM before async write (default: 10)
sysctl -w vm.dirty_ratio=40         # Increase for write performance
sysctl -w vm.dirty_background_ratio=5
sysctl -w vm.dirty_expire_centisecs=3000    # Age before writeback (30s)
sysctl -w vm.dirty_writeback_centisecs=500  # Writeback interval (5s)

# Transparent Huge Pages
cat /sys/kernel/mm/transparent_hugepage/enabled
echo never > /sys/kernel/mm/transparent_hugepage/enabled    # Disable (for databases)
echo always > /sys/kernel/mm/transparent_hugepage/enabled   # Enable

# OOM killer tuning
sysctl vm.panic_on_oom              # Panic vs kill on OOM
sysctl -w vm.panic_on_oom=0         # Kill process (don't panic)
sysctl vm.oom_kill_allocating_task  # Kill allocating task vs largest
sysctl -w vm.overcommit_memory=1    # Always overcommit (0=heuristic, 2=never)

# Page cache management
sysctl -w vm.vfs_cache_pressure=50  # Tendency to reclaim cache (default: 100)
# Lower = keep cache longer
# Higher = reclaim cache more aggressively

Filesystem and I/O Tuning

# File descriptor limits
sysctl fs.file-max                  # System-wide max open files
sysctl -w fs.file-max=2097152       # Increase limit

# Check current usage
cat /proc/sys/fs/file-nr
# openfiles  freelist  maximum

# Inotify limits (for file watching)
sysctl fs.inotify.max_user_watches
sysctl -w fs.inotify.max_user_watches=524288

# AIO limits
sysctl fs.aio-max-nr                # Max async I/O requests
sysctl -w fs.aio-max-nr=1048576

Kernel Tuning

# Kernel panic behaviour
sysctl kernel.panic                 # Seconds before reboot after panic
sysctl -w kernel.panic=10           # Reboot 10 seconds after panic

# Core dump settings
sysctl kernel.core_pattern
sysctl -w kernel.core_pattern=/var/crash/core.%e.%p.%t

# Process limits
sysctl kernel.pid_max               # Maximum PID
sysctl -w kernel.pid_max=4194304

# Shared memory
sysctl kernel.shmmax                # Max shared memory segment size
sysctl kernel.shmall                # Total shared memory pages
sysctl -w kernel.shmmax=68719476736 # 64 GB
sysctl -w kernel.shmall=4294967296  # 16 TB (pages)

Making sysctl Changes Persistent

# Temporary change (lost on reboot)
sysctl -w net.ipv4.tcp_fin_timeout=30

# Persistent change
echo "net.ipv4.tcp_fin_timeout = 30" >> /etc/sysctl.conf

# Or create file in sysctl.d
cat > /etc/sysctl.d/99-network.conf <<EOF
net.core.rmem_max = 134217728
net.core.wmem_max = 134217728
net.ipv4.tcp_rmem = 4096 87380 67108864
net.ipv4.tcp_wmem = 4096 65536 67108864
net.ipv4.tcp_congestion_control = bbr
EOF

# Apply changes
sysctl -p                           # Load from /etc/sysctl.conf
sysctl -p /etc/sysctl.d/99-network.conf    # Load specific file
sysctl --system                     # Load all configs

Quick Reference

Essential Performance Tools

Tool Purpose Key Usage
top / htop Real-time process monitor top -H (threads), htop -u user
vmstat Virtual memory stats vmstat 1 (per-second updates)
iostat I/O statistics iostat -x 1 (extended, per-second)
sar Historical system stats sar -u (CPU), sar -r (memory), sar -d (disk)
pidstat Per-process stats pidstat -u (CPU), pidstat -r (mem), pidstat -d (I/O)
mpstat Per-CPU statistics mpstat -P ALL 1
free Memory overview free -h
ss Socket statistics ss -tulpn (listening), ss -tan (all TCP)
perf CPU profiling perf top, perf record -g
iotop I/O by process iotop -o (active only)
iftop Network by connection iftop -i eth0

Key Performance Metrics

Metric Command Healthy Range Investigation Threshold
CPU load uptime < CPU count > 2x CPU count
CPU utilisation top, mpstat < 80% > 90% sustained
I/O wait vmstat, iostat < 5% > 10%
Disk utilisation iostat -x < 80% > 90% (%util column)
Memory available free -h > 10% total < 5% total
Swap usage free -h 0 or minimal Active swapping (si/so)
Context switches vmstat < 5000/s > 50000/s
Network errors ip -s link 0 Any errors/drops

Common Performance Commands

# Quick system overview
uptime && free -h && df -h

# CPU top consumers
ps aux --sort=-%cpu | head -20

# Memory top consumers
ps aux --sort=-%mem | head -20

# Disk I/O activity
iostat -x 1 5

# Network connections
ss -tan | wc -l

# System-wide profiling for 30 seconds
perf record -F 99 -a -g -- sleep 30 && perf report

# Check for resource exhaustion
ulimit -a                           # Process limits
sysctl fs.file-nr                   # System file descriptors
cat /proc/sys/kernel/pid_max        # Max processes

Common Issues and Solutions

Issue: High CPU Usage

Symptoms: System slow, load average high, CPU near 100%

Diagnosis:

# Identify CPU-intensive processes
top -o %CPU
ps aux --sort=-%cpu | head -10

# Check per-CPU usage
mpstat -P ALL 1 5

# Profile CPU usage
perf top
perf record -F 99 -a -g -- sleep 30
perf report

# Check for busy-waiting
pidstat -u 1 10

Solutions:

# Limit CPU with cgroups
echo "50000 100000" > /sys/fs/cgroup/myapp/cpu.max    # 50% CPU

# Adjust process priority
renice +10 -p PID                   # Lower priority

# Set CPU affinity
taskset -cp 0,1 PID                 # Pin to CPUs 0-1

# Scale down unnecessary services
systemctl stop service_name

Issue: Memory Exhaustion / OOM Kills

Symptoms: OOM killer messages, processes dying unexpectedly, swap usage high

Diagnosis:

# Check memory usage
free -h
vmstat 1 5

# Find memory hogs
ps aux --sort=-%mem | head -10
smem -t

# Check for OOM kills
dmesg | grep -i oom
journalctl -k | grep -i oom

# Monitor memory pressure
cat /proc/pressure/memory
watch -n 1 'cat /proc/meminfo | grep -E "MemAvailable|SwapFree"'

Solutions:

# Limit memory with cgroups
echo "2G" > /sys/fs/cgroup/myapp/memory.max

# Adjust OOM score to protect critical processes
echo -1000 > /proc/PID/oom_score_adj    # Never kill

# Reduce swappiness
sysctl -w vm.swappiness=10

# Increase swap space
dd if=/dev/zero of=/swapfile bs=1M count=4096
mkswap /swapfile
swapon /swapfile

# Restart memory-leaking applications
systemctl restart service_name

Issue: High I/O Wait

Symptoms: System slow, wa column high in top, high %util in iostat

Diagnosis:

# Check I/O statistics
iostat -x 1 5

# Identify I/O-intensive processes
iotop -o
pidstat -d 1 5

# Check disk queue
iostat -x | grep -E 'Device|sda' | grep -v loop

# Monitor specific device
iostat -x -d sda 1

# Check for slow queries (if database)
# MySQL:
mysql -e "SHOW PROCESSLIST"
# PostgreSQL:
psql -c "SELECT * FROM pg_stat_activity WHERE state = 'active'"

Solutions:

# Tune I/O scheduler
echo mq-deadline > /sys/block/sda/queue/scheduler

# Increase queue depth
echo 512 > /sys/block/sda/queue/nr_requests

# Limit I/O with cgroups
echo "8:0 rbps=104857600 wbps=52428800" > /sys/fs/cgroup/myapp/io.max

# Tune filesystem mount options
mount -o remount,noatime /mount/point

# Increase dirty page thresholds
sysctl -w vm.dirty_ratio=40
sysctl -w vm.dirty_background_ratio=10

# Optimise database configuration
# Add indices, tune buffer pools, etc.

Issue: Network Bottleneck

Symptoms: High packet loss, slow transfers, connection timeouts

Diagnosis:

# Check interface statistics
ip -s link
sar -n DEV 1 5

# Look for errors and drops
ip -s -s link show eth0

# Check connection states
ss -tan | awk '{print $1}' | sort | uniq -c

# Monitor bandwidth
iftop -i eth0
nload eth0

# Check TCP retransmissions
netstat -s | grep -i retrans
ss -ti | grep retrans

Solutions:

# Increase network buffers
sysctl -w net.core.rmem_max=134217728
sysctl -w net.core.wmem_max=134217728
sysctl -w net.ipv4.tcp_rmem="4096 87380 67108864"
sysctl -w net.ipv4.tcp_wmem="4096 65536 67108864"

# Tune TCP settings
sysctl -w net.ipv4.tcp_fin_timeout=30
sysctl -w net.ipv4.tcp_tw_reuse=1
sysctl -w net.core.somaxconn=4096

# Change congestion control
sysctl -w net.ipv4.tcp_congestion_control=bbr

# Increase connection tracking
sysctl -w net.netfilter.nf_conntrack_max=1048576

# Check NIC ring buffer
ethtool -g eth0
ethtool -G eth0 rx 4096 tx 4096        # Increase if supported

Issue: CPU Throttling in cgroups

Symptoms: Application slow despite low overall CPU, high throttled_usec in cpu.stat

Diagnosis:

# Check CPU throttling
cat /sys/fs/cgroup/myapp/cpu.stat | grep throttled
# nr_throttled   - Number of throttling periods
# throttled_usec - Total time throttled

# Monitor over time
watch -n 1 cat /sys/fs/cgroup/myapp/cpu.stat

Solutions:

# Increase CPU quota
echo "200000 100000" > /sys/fs/cgroup/myapp/cpu.max    # 2 CPUs

# Or remove limit entirely
echo "max 100000" > /sys/fs/cgroup/myapp/cpu.max

# For systemd services:
systemctl set-property nginx.service CPUQuota=200%

Issue: Memory Throttling in cgroups

Symptoms: Slow allocation, high latency, memory.events shows high or max hits

Diagnosis:

# Check memory events
cat /sys/fs/cgroup/myapp/memory.events

# Monitor memory usage vs limit
cat /sys/fs/cgroup/myapp/memory.current
cat /sys/fs/cgroup/myapp/memory.max

Solutions:

# Increase memory limit
echo "4G" > /sys/fs/cgroup/myapp/memory.max

# Increase soft limit to avoid throttling
echo "3G" > /sys/fs/cgroup/myapp/memory.high

# For systemd:
systemctl set-property nginx.service MemoryMax=4G
systemctl set-property nginx.service MemoryHigh=3G

Issue: High Context Switch Rate

Symptoms: System sluggish, high cs in vmstat, poor application performance

Diagnosis:

# Check context switch rate
vmstat 1 5                          # 'cs' column

# Per-process context switches
pidstat -w 1 5
pidstat -w -p PID 1

# Voluntary vs non-voluntary
pidstat -w -p PID | head -1 && pidstat -w -p PID 1 10

Solutions:

# Reduce number of threads
# Application-specific tuning

# Increase CPU quota if throttled
echo "max 100000" > /sys/fs/cgroup/myapp/cpu.max

# Disable power management
cpupower frequency-set -g performance

# Pin process to specific CPUs
taskset -cp 0-3 PID