The essentials

Quick reference

One focused task per row. Jump to the related section for complete, working examples.

UseSyntaxExamples
Show load and uptimeuptimeView examples
Sample system activityvmstat 1 5View examples
Count available CPUsnprocView examples
Sample every CPUmpstat -P ALL 1 5View examples
Sample process CPUpidstat -u -p ALL 1 5View examples
Summarize memoryfree --humanView examples
Sample process memorypidstat -r -p ALL 1 5View examples
Read memory pressurecat /proc/pressure/memoryView examples
Sample extended device I/Oiostat -xz 1 5View examples
Sample process I/Opidstat -d -p ALL 1 5View examples
Read I/O pressurecat /proc/pressure/ioView examples
Read historical CPU datasar -u -f /var/log/sa/sa12View examples
Sample queue and loadsar -q ALL 1 5View examples
Measure a bounded commandperf stat -- openssl speed -seconds 5 sha256View examples
List available eventsperf listView examples
Sample one processsudo perf record --call-graph dwarf --pid 4242 -- sleep \ 30View examples
Read a profileperf report --stdio --input=perf.dataView examples

Performance diagnosis is an evidence problem, not a tuning contest. Define the user-visible symptom and time window, check resource utilization, saturation, and errors, then narrow from host to cgroup, process, thread, and code path. Prefer bounded sampling and preserve workload context; profilers and high-frequency tracing add overhead, while cache-dropping and speculative sysctl changes destroy evidence and can worsen an incident.

Step by step

Detailed examples

01

Define a window and establish host-wide saturation signals

Load average counts runnable and uninterruptible tasks, not CPU percentage, so compare it with CPU capacity and blocked work. In vmstat, the first row is generally an average since boot; interpret later interval rows. Capture at least several samples spanning the symptom, timestamps, workload volume, host/cgroup limits, and recent deployment or failover events.

Take a bounded first-pass sample
date --iso-8601=seconds
uptime
nproc
vmstat 1 5
cat /proc/pressure/cpu
Back to quick reference ↑
02

Separate CPU utilization, runnable contention, and imbalance

High user or system time means active CPU work; a growing run queue and CPU pressure indicate tasks waiting for execution. Low aggregate utilization can hide one saturated CPU, affinity constraint, or single-thread bottleneck. Compare mpstat per-CPU samples with pidstat attribution and cgroup quotas before assuming more host CPUs will help.

Correlate CPUs with consuming processes
mpstat -P ALL 1 5
pidstat -u -p ALL 1 5
cat /proc/pressure/cpu
Back to quick reference ↑
03

Use available memory, reclaim, faults, swap, and PSI together

Linux intentionally uses free memory for cache, so a small free value is not itself a leak. MemAvailable estimates allocatable capacity; vmstat si/so reveals active swap traffic, pidstat shows process faults and RSS, and PSI quantifies time lost to memory contention. Full pressure means all non-idle work is stalled and is often more urgent than a capacity percentage.

Measure memory behavior without clearing caches
free --human
vmstat 1 5
pidstat -r -p ALL 1 5
cat /proc/pressure/memory
Back to quick reference ↑
04

Correlate device latency with workload I/O and stalls

iostat -x exposes throughput, queueing, latency, and busy time, but field names and semantics vary by sysstat version and device type. High utilization is not universally saturation on parallel devices. Compare multiple intervals, I/O pressure, filesystem/storage-layer errors, and pidstat attribution; device-mapper, RAID, network storage, and containers can separate the apparent issuer from the physical bottleneck.

Triangulate a storage slowdown
iostat -xz 1 5
pidstat -d -p ALL 1 5
cat /proc/pressure/io
journalctl --dmesg --boot --priority=warning --no-pager
Back to quick reference ↑
05

Use sysstat archives to reconstruct transient incidents

sar reads periodic samples collected by sysstat when collection and retention are enabled. Confirm the archive date, host timezone, sampling cadence, restart markers, and counter units. Never assume a path naming convention across distributions; /var/log/sa/saDD is common. Export or preserve relevant archives before retention rotates them.

Inspect a known archive and current queue state
sar -u -f /var/log/sa/sa12
sar -r -f /var/log/sa/sa12
sar -q ALL 1 5
Back to quick reference ↑
06

Use perf stat to test a specific CPU hypothesis

perf stat counts events over a bounded workload and can report elapsed time, cycles, instructions, branches, faults, and derived metrics. Available events, multiplexing, virtualization accuracy, hybrid processors, and kernel perf_event_paranoid policy affect results. Warm up repeatable workloads, compare multiple runs, and record kernel, CPU, affinity, frequency policy, and event availability.

Measure a finite CPU workload
perf list
perf stat --repeat 3 -- openssl speed -seconds 5 sha256
# Review event coverage and scaling percentages before comparing runs.
Back to quick reference ↑
07

Scope sampling profiles by process and duration

perf record samples execution and can add measurable CPU, storage, and unwind overhead. Obtain change approval, attach only to the target PID, keep the duration short, and protect perf.data because stacks and symbol names may expose sensitive implementation details. DWARF call graphs cost more than frame-pointer unwinding; validate unwind quality and delete or archive data according to incident policy.

Capture and inspect a 30-second approved profile
sudo perf record --call-graph dwarf --pid 4242 -- sleep 30
perf report --stdio --input=perf.data
# Treat perf.data as sensitive operational evidence.
Back to quick reference ↑

Sources and further reading

References

Authoritative documentation used to verify and expand this cheat sheet.

  1. Linux Kernel documentationPSI — Pressure Stall Informationkernel.org
  2. Linux perf tools via Linux man-pagesperf-stat(1)man7.org
  3. procps-ng Project via Linux man-pagesvmstat(8)man7.org
  4. sysstat Project via Linux man-pagespidstat(1)man7.org
  5. sysstat Project via Linux man-pagesiostat(1)man7.org
  6. sysstat Project via Linux man-pagessar(1)man7.org

Help us improve

Found a typo or missing example?

Tell us what would make this cheat sheet clearer, more complete, or more useful.

Share feedback