The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Read a monotonic performance clock | started = time.perf_counter() | View examples |
| Read integer nanoseconds | started_ns = time.perf_counter_ns() | View examples |
| Record interpreter identity | identity = (sys.implementation.name, platform.python_version()) | View examples |
| Time a callable | seconds = timeit.timeit(callable_under_test, number=1000) | View examples |
| Collect repeated timings | samples = timeit.repeat(stmt, setup=setup, repeat=7, number=1000) | View examples |
| Run the timeit CLI | python -m timeit -s 'data = list(range(100))' 'sum(data)' | View examples |
| Profile a callable | profile.runcall(main) | View examples |
| Profile a module to a file | python -m cProfile -o profile.pstats -m package.module | View examples |
| Sort by cumulative time | stats.sort_stats(pstats.SortKey.CUMULATIVE).print_stats(20) | View examples |
| Start allocation tracing | tracemalloc.start(10) | View examples |
| Capture allocations | snapshot = tracemalloc.take_snapshot() | View examples |
| Compare snapshots | diff = after.compare_to(before, 'lineno') | View examples |
| Install a trace hook | previous = sys.gettrace(); sys.settrace(trace_function) | View examples |
| Restore a trace hook | sys.settrace(previous) | View examples |
| Trace threading workers | threading.settrace(trace_function) | View examples |
| Enable fault tracebacks | faulthandler.enable() | View examples |
| Inspect current thread frames | frames = sys._current_frames() | View examples |
| Summarize the median | middle = statistics.median(samples) | View examples |
| Compute quartile cut points | quartiles = statistics.quantiles(samples, n=4, method='inclusive') | View examples |
Performance work starts with a reproducible workload and a question. Benchmark elapsed behavior with controlled repetitions, profile to find where time accumulates, trace allocations to find where memory originates, and use tracing hooks only when their semantic detail justifies their overhead. Measurements are evidence from one environment, not timeless constants.
Step by step
Detailed examples
Define the workload before measuring
Warm caches when that matches production, isolate setup from the operation under test, and record Python version, build, platform, dependency versions, input shape, and statistical method. Compare equivalent outputs before comparing speed. CPU frequency scaling, background work, garbage collection, and virtualized hosts can dominate microbenchmarks.
def workload(values: list[int]) -> int:
return sum(value * value for value in values)
result = workload([1, 2, 3, 4])
assert result == 30
print("verified", result) verified 30Use timeit for focused operations
timeit constructs a small repeatable harness and uses perf_counter by default. Put imports and data creation in setup when they are not part of the operation, use callable statements to avoid string-scope surprises, and inspect multiple repeats rather than trusting one run. The minimum often estimates least-contended execution; report the full distribution when variability matters.
import timeit
calls = 0
def operation() -> None:
global calls
calls += 1
timeit.timeit(operation, number=5)
print(calls) 5Capture deterministic call profiles
cProfile records call counts and cumulative/internal time with relatively low overhead. Profile a representative entry point, save raw stats for repeatable analysis, and use pstats sorting and restrictions to focus output. Instrumentation perturbs execution, so do not interpret its timing as an unbiased benchmark.
import cProfile
import pstats
def double(value: int) -> int:
return value * 2
profile = cProfile.Profile()
profile.enable()
for value in range(3):
double(value)
profile.disable()
stats = pstats.Stats(profile)
calls = sum(item[0] for key, item in stats.stats.items() if key[2] == "double")
print(calls) 3Attribute allocations with tracemalloc
Start tracemalloc before the allocations of interest and choose enough frames to recover useful call paths. Snapshots compare currently traced allocations, not resident set size, native allocator use, or every extension-module allocation. Filters remove harness noise; stop tracing after capture because retaining stack traces costs time and memory.
import tracemalloc
tracemalloc.start()
before, _ = tracemalloc.get_traced_memory()
payload = [str(value) for value in range(100)]
after, peak = tracemalloc.get_traced_memory()
print(len(payload))
print(after >= before)
print(peak >= after)
tracemalloc.stop() 100
True
TrueUse trace hooks for semantic events
sys.settrace receives call, line, return, exception, and optional opcode events for the current thread. The hook must return the local trace function for scopes it wants to continue tracing. Hooks are thread-specific, impose high overhead, and may expose sensitive local values; install them narrowly and restore the previous function in finally.
import sys
calls = 0
def target() -> int:
return 42
def tracer(frame, event, arg):
global calls
if event == "call" and frame.f_code is target.__code__:
calls += 1
return tracer
previous = sys.gettrace()
try:
sys.settrace(tracer)
target()
target()
finally:
sys.settrace(previous)
print(calls) 2Prefer low-overhead evidence in production
Statistical profilers periodically sample stacks and can reveal dominant paths with less distortion than per-event tracing. Sample long enough to capture representative traffic and segment results by workload. Native frames, subprocesses, containers, permissions, and optimized interpreter builds change what a profiler can observe; validate tool support in a staging environment.
samples = ["serve>parse", "serve>parse", "serve>write", "idle"]
counts: dict[str, int] = {}
for stack in samples:
counts[stack] = counts.get(stack, 0) + 1
for stack in sorted(counts):
print(stack, counts[stack]) idle 1
serve>parse 2
serve>write 1Report uncertainty and prevent benchmark traps
State the tested hypothesis, environment, workload, sample count, warm-up, summary statistic, dispersion, and raw data. Avoid rounding away meaningful differences, but do not claim precision the clock and system cannot support. Run A/B variants in alternating or randomized order, inspect regressions on stable hardware, and optimize only after correctness tests show identical behavior.
import statistics
baseline = [10.0, 10.2, 9.8]
candidate = [8.0, 8.1, 7.9]
ratio = statistics.median(candidate) / statistics.median(baseline)
print(f"candidate/baseline={ratio:.2f}") candidate/baseline=0.80Local code tester
Summarize benchmark samples
Compare two deterministic sample sets with medians and a normalized ratio.
Press Run to load Python locally.
Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
- Python Software FoundationThe Python Profilersdocs.python.org
- Python Software Foundationtimeit — Measure execution time of small code snippetsdocs.python.org
- Python Software Foundationtracemalloc — Trace memory allocationsdocs.python.org
- Python Software Foundationsys — settrace and setprofiledocs.python.org
- Python Software Foundationfaulthandler — Dump the Python tracebackdocs.python.org
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



