The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Show load and uptime | uptime | View examples |
| Sample system activity | vmstat 1 5 | View examples |
| Count available CPUs | nproc | View examples |
| Sample every CPU | mpstat -P ALL 1 5 | View examples |
| Sample process CPU | pidstat -u -p ALL 1 5 | View examples |
| Summarize memory | free --human | View examples |
| Sample process memory | pidstat -r -p ALL 1 5 | View examples |
| Read memory pressure | cat /proc/pressure/memory | View examples |
| Sample extended device I/O | iostat -xz 1 5 | View examples |
| Sample process I/O | pidstat -d -p ALL 1 5 | View examples |
| Read I/O pressure | cat /proc/pressure/io | View examples |
| Read historical CPU data | sar -u -f /var/log/sa/sa12 | View examples |
| Sample queue and load | sar -q ALL 1 5 | View examples |
| Measure a bounded command | perf stat -- openssl speed -seconds 5 sha256 | View examples |
| List available events | perf list | View examples |
| Sample one process | sudo perf record --call-graph dwarf --pid 4242 -- sleep \
30 | View examples |
| Read a profile | perf report --stdio --input=perf.data | View examples |
Performance diagnosis is an evidence problem, not a tuning contest. Define the user-visible symptom and time window, check resource utilization, saturation, and errors, then narrow from host to cgroup, process, thread, and code path. Prefer bounded sampling and preserve workload context; profilers and high-frequency tracing add overhead, while cache-dropping and speculative sysctl changes destroy evidence and can worsen an incident.
Step by step
Detailed examples
Define a window and establish host-wide saturation signals
Load average counts runnable and uninterruptible tasks, not CPU percentage, so compare it with CPU capacity and blocked work. In vmstat, the first row is generally an average since boot; interpret later interval rows. Capture at least several samples spanning the symptom, timestamps, workload volume, host/cgroup limits, and recent deployment or failover events.
date --iso-8601=seconds
uptime
nproc
vmstat 1 5
cat /proc/pressure/cpu Separate CPU utilization, runnable contention, and imbalance
High user or system time means active CPU work; a growing run queue and CPU pressure indicate tasks waiting for execution. Low aggregate utilization can hide one saturated CPU, affinity constraint, or single-thread bottleneck. Compare mpstat per-CPU samples with pidstat attribution and cgroup quotas before assuming more host CPUs will help.
mpstat -P ALL 1 5
pidstat -u -p ALL 1 5
cat /proc/pressure/cpu Use available memory, reclaim, faults, swap, and PSI together
Linux intentionally uses free memory for cache, so a small free value is not itself a leak. MemAvailable estimates allocatable capacity; vmstat si/so reveals active swap traffic, pidstat shows process faults and RSS, and PSI quantifies time lost to memory contention. Full pressure means all non-idle work is stalled and is often more urgent than a capacity percentage.
free --human
vmstat 1 5
pidstat -r -p ALL 1 5
cat /proc/pressure/memory Correlate device latency with workload I/O and stalls
iostat -x exposes throughput, queueing, latency, and busy time, but field names and semantics vary by sysstat version and device type. High utilization is not universally saturation on parallel devices. Compare multiple intervals, I/O pressure, filesystem/storage-layer errors, and pidstat attribution; device-mapper, RAID, network storage, and containers can separate the apparent issuer from the physical bottleneck.
iostat -xz 1 5
pidstat -d -p ALL 1 5
cat /proc/pressure/io
journalctl --dmesg --boot --priority=warning --no-pager Use sysstat archives to reconstruct transient incidents
sar reads periodic samples collected by sysstat when collection and retention are enabled. Confirm the archive date, host timezone, sampling cadence, restart markers, and counter units. Never assume a path naming convention across distributions; /var/log/sa/saDD is common. Export or preserve relevant archives before retention rotates them.
sar -u -f /var/log/sa/sa12
sar -r -f /var/log/sa/sa12
sar -q ALL 1 5 Use perf stat to test a specific CPU hypothesis
perf stat counts events over a bounded workload and can report elapsed time, cycles, instructions, branches, faults, and derived metrics. Available events, multiplexing, virtualization accuracy, hybrid processors, and kernel perf_event_paranoid policy affect results. Warm up repeatable workloads, compare multiple runs, and record kernel, CPU, affinity, frequency policy, and event availability.
perf list
perf stat --repeat 3 -- openssl speed -seconds 5 sha256
# Review event coverage and scaling percentages before comparing runs. Scope sampling profiles by process and duration
perf record samples execution and can add measurable CPU, storage, and unwind overhead. Obtain change approval, attach only to the target PID, keep the duration short, and protect perf.data because stacks and symbol names may expose sensitive implementation details. DWARF call graphs cost more than frame-pointer unwinding; validate unwind quality and delete or archive data according to incident policy.
sudo perf record --call-graph dwarf --pid 4242 -- sleep 30
perf report --stdio --input=perf.data
# Treat perf.data as sensitive operational evidence. Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
- Linux Kernel documentationPSI — Pressure Stall Informationkernel.org
- Linux perf tools via Linux man-pagesperf-stat(1)man7.org
- procps-ng Project via Linux man-pagesvmstat(8)man7.org
- sysstat Project via Linux man-pagespidstat(1)man7.org
- sysstat Project via Linux man-pagesiostat(1)man7.org
- sysstat Project via Linux man-pagessar(1)man7.org
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



