The essentials

Quick reference

One focused task per row. Jump to the related section for complete, working examples.

UseSyntaxExamples
Read a monotonic performance clockstarted = time.perf_counter()View examples
Read integer nanosecondsstarted_ns = time.perf_counter_ns()View examples
Record interpreter identityidentity = (sys.implementation.name, platform.python_version())View examples
Time a callableseconds = timeit.timeit(callable_under_test, number=1000)View examples
Collect repeated timingssamples = timeit.repeat(stmt, setup=setup, repeat=7, number=1000)View examples
Run the timeit CLIpython -m timeit -s 'data = list(range(100))' 'sum(data)'View examples
Profile a callableprofile.runcall(main)View examples
Profile a module to a filepython -m cProfile -o profile.pstats -m package.moduleView examples
Sort by cumulative timestats.sort_stats(pstats.SortKey.CUMULATIVE).print_stats(20)View examples
Start allocation tracingtracemalloc.start(10)View examples
Capture allocationssnapshot = tracemalloc.take_snapshot()View examples
Compare snapshotsdiff = after.compare_to(before, 'lineno')View examples
Install a trace hookprevious = sys.gettrace(); sys.settrace(trace_function)View examples
Restore a trace hooksys.settrace(previous)View examples
Trace threading workersthreading.settrace(trace_function)View examples
Enable fault tracebacksfaulthandler.enable()View examples
Inspect current thread framesframes = sys._current_frames()View examples
Summarize the medianmiddle = statistics.median(samples)View examples
Compute quartile cut pointsquartiles = statistics.quantiles(samples, n=4, method='inclusive')View examples

Performance work starts with a reproducible workload and a question. Benchmark elapsed behavior with controlled repetitions, profile to find where time accumulates, trace allocations to find where memory originates, and use tracing hooks only when their semantic detail justifies their overhead. Measurements are evidence from one environment, not timeless constants.

Step by step

Detailed examples

01

Define the workload before measuring

Warm caches when that matches production, isolate setup from the operation under test, and record Python version, build, platform, dependency versions, input shape, and statistical method. Compare equivalent outputs before comparing speed. CPU frequency scaling, background work, garbage collection, and virtualized hosts can dominate microbenchmarks.

Verify work before recording a duration
def workload(values: list[int]) -> int:
    return sum(value * value for value in values)

result = workload([1, 2, 3, 4])
assert result == 30
print("verified", result)
Output
verified 30
Back to quick reference ↑
02

Use timeit for focused operations

timeit constructs a small repeatable harness and uses perf_counter by default. Put imports and data creation in setup when they are not part of the operation, use callable statements to avoid string-scope surprises, and inspect multiple repeats rather than trusting one run. The minimum often estimates least-contended execution; report the full distribution when variability matters.

Confirm the timeit harness call count
import timeit

calls = 0
def operation() -> None:
    global calls
    calls += 1

timeit.timeit(operation, number=5)
print(calls)
Output
5
Back to quick reference ↑
03

Capture deterministic call profiles

cProfile records call counts and cumulative/internal time with relatively low overhead. Profile a representative entry point, save raw stats for repeatable analysis, and use pstats sorting and restrictions to focus output. Instrumentation perturbs execution, so do not interpret its timing as an unbiased benchmark.

Inspect stable profiler call counts
import cProfile
import pstats

def double(value: int) -> int:
    return value * 2

profile = cProfile.Profile()
profile.enable()
for value in range(3):
    double(value)
profile.disable()
stats = pstats.Stats(profile)
calls = sum(item[0] for key, item in stats.stats.items() if key[2] == "double")
print(calls)
Output
3
Back to quick reference ↑
04

Attribute allocations with tracemalloc

Start tracemalloc before the allocations of interest and choose enough frames to recover useful call paths. Snapshots compare currently traced allocations, not resident set size, native allocator use, or every extension-module allocation. Filters remove harness noise; stop tracing after capture because retaining stack traces costs time and memory.

Use traced-memory counters without fixed byte claims
import tracemalloc

tracemalloc.start()
before, _ = tracemalloc.get_traced_memory()
payload = [str(value) for value in range(100)]
after, peak = tracemalloc.get_traced_memory()
print(len(payload))
print(after >= before)
print(peak >= after)
tracemalloc.stop()
Output
100
True
True
Back to quick reference ↑
05

Use trace hooks for semantic events

sys.settrace receives call, line, return, exception, and optional opcode events for the current thread. The hook must return the local trace function for scopes it wants to continue tracing. Hooks are thread-specific, impose high overhead, and may expose sensitive local values; install them narrowly and restore the previous function in finally.

Count calls to one target function
import sys

calls = 0
def target() -> int:
    return 42

def tracer(frame, event, arg):
    global calls
    if event == "call" and frame.f_code is target.__code__:
        calls += 1
    return tracer

previous = sys.gettrace()
try:
    sys.settrace(tracer)
    target()
    target()
finally:
    sys.settrace(previous)
print(calls)
Output
2
Back to quick reference ↑
06

Prefer low-overhead evidence in production

Statistical profilers periodically sample stacks and can reveal dominant paths with less distortion than per-event tracing. Sample long enough to capture representative traffic and segment results by workload. Native frames, subprocesses, containers, permissions, and optimized interpreter builds change what a profiler can observe; validate tool support in a staging environment.

Summarize a synthetic stack sample
samples = ["serve>parse", "serve>parse", "serve>write", "idle"]
counts: dict[str, int] = {}
for stack in samples:
    counts[stack] = counts.get(stack, 0) + 1
for stack in sorted(counts):
    print(stack, counts[stack])
Output
idle 1
serve>parse 2
serve>write 1
Back to quick reference ↑
07

Report uncertainty and prevent benchmark traps

State the tested hypothesis, environment, workload, sample count, warm-up, summary statistic, dispersion, and raw data. Avoid rounding away meaningful differences, but do not claim precision the clock and system cannot support. Run A/B variants in alternating or randomized order, inspect regressions on stable hardware, and optimize only after correctness tests show identical behavior.

Report a deterministic normalized comparison
import statistics

baseline = [10.0, 10.2, 9.8]
candidate = [8.0, 8.1, 7.9]
ratio = statistics.median(candidate) / statistics.median(baseline)
print(f"candidate/baseline={ratio:.2f}")
Output
candidate/baseline=0.80
Back to quick reference ↑

Local code tester

Summarize benchmark samples

Compare two deterministic sample sets with medians and a normalized ratio.

Runs in your browser
Output
Press Run to load Python locally.

Sources and further reading

References

Authoritative documentation used to verify and expand this cheat sheet.

  1. Python Software FoundationThe Python Profilersdocs.python.org
  2. Python Software Foundationtimeit — Measure execution time of small code snippetsdocs.python.org
  3. Python Software Foundationtracemalloc — Trace memory allocationsdocs.python.org
  4. Python Software Foundationsys — settrace and setprofiledocs.python.org
  5. Python Software Foundationfaulthandler — Dump the Python tracebackdocs.python.org

Help us improve

Found a typo or missing example?

Tell us what would make this cheat sheet clearer, more complete, or more useful.

Share feedback