vperf wraps Linux perf (and macOS's sample) and answers the questions a
profile is for: where did the time go, why did the cores stall, which cache
levels did the memory traffic hit, and what were the threads doing when they
weren't on a CPU. One command, one self-contained HTML report. AMD and Intel,
Python ≥ 3.10, zero runtime dependencies.
# 1. allow user-space profiling (one-time, needs sudo)
sudo sysctl kernel.perf_event_paranoid=1
# 2. check what this machine supports
vperf doctor
# 3. profile a program
vperf run -- ./yourapp input.bin
# 4. open the report
xdg-open .vperf/run_*/report.htmlvperf run prints a summary to the terminal and writes report.html into a
fresh .vperf/run_<timestamp>/ directory. The report is a single HTML file that
works offline in any browser — no CDN, no server, nothing fetched from the
network.
No install step either — with a checkout, the system Python is enough:
python3 -m vperf run -- ./yourapp input.binFor development, uv venv && uv pip install -e . gives you the vperf console
script. See Development.
- Overview — CPU time, effective utilization, IPC/CPI, branch mispredicts, cache and TLB miss rates, and a Backend/Frontend Bound breakdown.
- Hotspots — self and inclusive time per function.
- Flame graph — click a frame to zoom into that branch.
- Call tree — the same stacks, collapsible.
- Memory access — AMD IBS or Intel PEBS samples by cache level (L1 / L3 / DRAM) with latency bands, the top symbols, and a timeline over the run.
- Wait / off-CPU — per thread, On-CPU seconds next to Sleep, Blocked/IO and Runnable seconds, with a "where the time went" bar.
- Timelines — CPU utilization, resident memory (with the run's peak) and
frequency, across the whole run. The memory and frequency curves are placed on
the sample timeline by a clock origin each sampler records, and that clock is
not one userspace is guaranteed to be reading: inside a Linux time namespace
CLOCK_MONOTONICis shifted by the namespace's offset, and perf's own clock can run ahead ofCLOCK_BOOTTIMEby a drift that grows with uptime — measured here as 0.38 s at 2.3 h of uptime and 3.9 s at 11.8 h, about +1.6 s a day. So vperf measures that offset per profile, from where the recording started and ended against where its samples fall, and uses it to place the curves — unless the measurement demonstrably puts the readings outside the measured window, in which case the curve is anchored to the first sample instead, because on such a host much of the gap is the delay before perf's first sample rather than a clock difference. Either way the curve is drawn only where it was measured: a stretch with no reading is left blank and labelled, never continued, since a held value is a measurement nobody took. Left uncorrected both curves draw empty while their data sits in the file, which is invisible rather than wrong-looking. - Scoping — a thread selector (with an optional group-by-name mode) and a time selection on the chart narrow every view that has the data for it.
Hardware a machine doesn't expose reads n/a rather than being guessed at.
vperf doctor prints exactly what your machine supports.
| Area | AMD (Zen) | Intel |
|---|---|---|
| Memory-access sampling | ibs_op (sampling period, default 100003) |
mem-loads/mem-stores PEBS (load-latency threshold, default --ldlat 30) |
| LLC classification | ls_any_fills_from_sys.* — DRAM/MMIO fills per L3 lookup |
LLC-loads/LLC-load-misses — L1 misses that reached L3 |
| L1 / L2 miss counts | ls_any_fills_from_sys.all, l2_cache_req_stat.ic_dc_miss_in_l2 |
not exposed; L1-dcache-load-misses drives the L1D miss rate only |
| FP / vectorization | fp_ret_sse_avx_ops.*, fp_ops_retired_by_width.* |
not exposed; Overview omits the FP rows |
| Branch-mispredict penalty (Bad Speculation model) | 13 cyc | 15 cyc |
Only the AMD-only events are vendor-gated, so Intel never sees <not counted>
noise. Everything else is collected and reported identically on both.
| Flag | Use it when |
|---|---|
-o, --outdir NAME |
you want a named profile directory — vperf diff compares two |
-f, --freq HZ (199) |
the default sample rate misses short phases |
--callgraph dwarf |
the target has good DWARF; fp (the default) needs no debug info but requires the target to preserve frame pointers |
--startup-grace SEC (0.15) |
the target spawns its thread pool during startup — per-thread counters only see the threads alive when counting attaches |
--mem-period N (100003) |
a long multi-threaded run produces millions of AMD IBS samples; raise the period to thin them |
--no-inline |
a huge C++ binary spends most of the run expanding DWARF inlines (measured: a 4.9 GB ClickHouse debug build, ~80 s per perf call with inlines, ~2 s without) |
--no-stat / --no-wait / --no-rss |
skip a pass you don't need — each one removes its panel from the report |
--mem-time-quantum MS |
finer or coarser Memory-tab slices (default ~100 over the run, clamped 25 ms–1 s) |
--keep-perf-data |
you want the raw perf.data left in the profile directory — it is removed by default once the dumps have been taken from it, and it is ~100 MB per second of a wide target |
--startup-grace is worth calibrating against your target: measured on
clickhouse-local, the thread pool goes 1 thread at 0 ms, 3 at 11 ms, 25 at
42 ms and 43 at 93 ms. Counters only start once the target resumes, so a longer
grace costs no measurement accuracy — but a target that exits inside the window
is reported rather than profiled, so keep the value below the shortest run you
care about.
vperf attach -p 1234 --duration 10 # profile a running process (needs CAP_PERFMON)
vperf report .vperf/run_20260824_021912 # regenerate report.html from a saved profile
vperf diff .vperf/before .vperf/after # compare two profilesvperf report replays the text artifacts in the profile directory, so it works on
a directory the raw recording has been removed from. Pass --keep-perf-data when
you want to re-derive the dumps yourself — perf script with the DWARF inlines
left in, perf report, perf archive, or a different tool entirely.
A single profile is noisy, so vperf cycle repeats a target N times and writes a
TSV matrix of metrics — one row per run — which ministat (from the BSD
toolkit) turns into a verdict at 95% confidence:
vperf cycle -n 30 -- ./myapp > base.tsv
vperf cycle -n 30 -- ./myapp-fixed > fixed.tsv
awk -F'\t' 'NR>1{print $4}' base.tsv > base_ipc.txt # column 4 is IPC
awk -F'\t' 'NR>1{print $4}' fixed.tsv > fixed_ipc.txt
ministat base_ipc.txt fixed_ipc.txtOptions: -n measured runs (default 30), --warmup discarded runs (default 1),
-j parallel runs (default = half your logical CPUs, pinned round-robin to
distinct physical cores so SMT siblings do not contend), --metrics CSV to pick
columns, --tsv FILE instead of stdout. Progress and the per-metric summary go to
stderr, so > file.tsv stays clean.
macOS has no perf and none of the PMU / tracepoint surface the counting passes
need, so vperf there runs a sample-based backend: hotspots, flame graph,
call tree, RSS and a utilization timeline, all in the same report; every
counter-driven metric reads n/a. Requires Xcode Command Line Tools (for
/usr/bin/sample). vperf doctor reports what is available. The counter flags
(--callgraph, --no-inline, --mem-period, --freq) are ignored there, as is
--keep-perf-data — there is no perf.data to retire — and vperf cycle is
unavailable.
- The Flame Graph and Call Tree show user-space frames only; kernel frames
become a synthetic
[kernel boundary]leaf that keeps their sample weight. Hotspots and metrics keep the full data. - Frame pointers are the default unwinder and need no debug info, but the target
must preserve them. Missing frame pointers degrade stacks — use
--callgraph dwarffor optimized binaries with good unwind data. Fully stripped binaries give address-level samples only. - Counters share PMU registers, so ratios across different multiplexed groups carry some noise.
- Backend/Frontend Bound are stall ratios, not TMA slot fractions: perf's
tma_*metrics need Intel'sslotsPMU, whichvperfdoes not request. Bad Speculation and Retiring are a model; the assumed recovery penalty is printed next to the result. - Per-thread counters cover only the threads alive when counting attaches (see
--startup-grace). Threads created afterwards are unavailable rather than estimated.
- A capability probe runs tiny throwaway profiles to find which events,
-Mmetrics and precise cycles event this machine supports. - The target is frozen, then one synchronized
perf stat --per-thread+perf recordsession collects counters, CPU samples and memory-access samples from the same workload lifetime, with the scheduler tracepoints co-joined into the same recording. - After the target exits,
perf scriptandperf mem reportreadperf.datain parallel, then the dumps are parsed in pure Python. The recording is removed once both are done —vperf reportnever needs it — unless--keep-perf-datawas passed. - Metrics are derived and
report.htmlis rendered as a single file.
AGENTS.md documents the pipeline, the artifacts and the internal contracts.
pytest and ruff are the only things a virtualenv is for; vperf itself runs
uninstalled.
uv venv && uv pip install -e . pytest ruff
source .venv/bin/activate
vperf doctor # verify this host can be profiled
pytest tests/ -q # unit + integration (needs perf access)
ruff check vperf/ tests/
# Note: run the suite WITHOUT pytest-xdist/-n. The integration tests assert
# exact PMU counter relationships; concurrent profiling sessions multiplex
# the hardware counters and break those assertions.examples/ holds C++ workloads with opposite, well-understood hardware
signatures, used by the integration tests:
make -C examples
vperf run -- examples/bin/simd_levels_avx 2000000 # AVX2, L1-resident -> IPC ~2.9
vperf run -- examples/bin/membound 268435456 1.5 # 256 MiB chase -> IPC ~0.14AGENTS.md is the contributor guide — module map, data flow, the
report UI contract, internal invariants and conventions. Read it before
changing anything under vperf/.
bench/clickbench_profiles.sh profiles every ClickBench query against
clickhouse-local and duckdb, one profile directory per query and engine, and
keeps the statement it ran in queries.sql next to the report:
bench/clear-caches.py # evict the dataset, report residency (no root)
bench/clickbench_profiles.sh --engine both --from 13 --to 13Each query runs exactly once — nothing repeated, retimed or retried — and the
engines run sequentially, because concurrent profiling sessions multiplex the
hardware counters. bench/clear-caches.py makes a cold run known rather than
assumed (posix_fadvise(POSIX_FADV_DONTNEED), verified with mincore), which
is how the I/O half of a query becomes visible: the same Q13 on the same machine
spent 1.3 s in Blocked/IO cold against 0.3 s warm, 12.3 s of CPU against 22.7 s.
--mode server runs the same 43 queries against a MergeTree hits table in a
clickhouse-server rather than clickhouse-local reading the parquet, and
profiles each one by attaching vperf to the server:
bench/clickbench_profiles.sh --mode server --load # load once, then profile all
bench/clickbench_profiles.sh --mode server --from 18 --to 18 # one query, table reused
bench/clickbench_profiles.sh --mode server --optimize-final # merge every part first
bench/clickbench_profiles.sh --mode server --private-server # own server, own data dirLoading takes no privileges at all. ClickHouse confines the file() function to
its user_files_path (/var/lib/clickhouse/user_files/, inside a 700
clickhouse-owned /var/lib), so a server cannot read a parquet that sits
anywhere else — ClickBench's own load works around that with a root-owned
symlink. The driver streams the file's bytes in instead
(INSERT INTO hits FORMAT Parquet, parsed by the server with parallel parsing),
so nothing is installed, nothing is added to a group, and no sudo is involved.
The dataset's columns are checked against the schema before the first byte is
sent, because the Parquet reader matches columns by name and would otherwise
fill a missing one with defaults.
Credentials come from ~/.clickhouse-client/config.xml (<host>, <port>,
<user>, <password>) — the driver has no auth flags of its own, so the
server's credentials stay in one place you control; --private-server is the
exception it has to spell out, and passes its own --port.
Loading is opt-in: --load is the only thing that loads. Everything else
reuses the hits table the server already has, and errors naming --load if it
has none. Loading is 100M rows and about 200 s, the table is reusable
afterwards, and it is never the thing being measured — a sweep that quietly
refilled it every run looked identical to one that never did. --load drops and
refills, so it is also how you pick up a changed schema. Loading 100M rows into a
table with fsync_after_insert = 1 takes minutes, and the log carries the table's
size as it fills. --optimize / --optimize-final are opt-in because ClickBench
does not optimize, and because the schema has no PARTITION BY: a FINAL merges
every part of the whole table.
With --private-server the table outlives the run: its data directory
(.vperf/_private-clickhouse) is reused rather than recreated, so the second
sweep finds the loaded table already there. rm -rf it to start over, or to
reclaim the parts a killed server leaves un-merged — a server stopped the moment
the sweep ends never gets to merge them, and that measured 30 GB on disk for a
9.1 GiB table.
Two things to know about the reports:
- A profile is the whole server process, not the query. A running
ClickHouse holds hundreds of background threads (164
ThreadPool, 16MergeMutate, 16Fetchand theBg*pools on a stock 16-core box), so the Overview counters and the utilization curve cover all of them — group or scope to the query's threads in the Threads tab to read the query.--private-serverstarts a dedicated server instead, which keeps those threads out of the profile. - The private server is the forked child, not the one you launched.
clickhouse serverstarts a lightweight supervisor and forks the real server into a second process, so the pid the driver is handed has 7 threads and noThreadPoolworkers while the server has 318. The driver profiles the process that owns the port the query goes to — profiling the supervisor instead yields a near-empty report (measured: 0.03 s of CPU against 7.45 s, 62 samples against 1.25K) that looks like a broken profiler rather than a wrong pid. - Each profile covers exactly the query. The runtime is not knowable in
advance (Q00 is ~0.1 s, Q35 ~90 s), so a fixed window would either truncate the
query or pad it with idle server; vperf ends the profile the moment the client
returns.
--max-duration(300 s) is only the ceiling.
Three things differ from a single-process target, all because sampling cost is per thread and a server has hundreds of them:
- Lower default rates.
FREQdefaults to 99 Hz here instead of 499, andMEM_PERIODto 4000003 instead of 1000003: aclickhouse-serverruns 359 threads on a stock 16-core box (164ThreadPool, 16MergeMutate, 16Fetch, theBg*pools), and IBS samples every thread that retires cycles, including background merges. Override withFREQ=499 MEM_PERIOD=1000003. - No scheduler tracepoints (
--waitopts back in): a server switches hundreds of background threads and their off-CPU time is not what a query report is about. - No freeze. vperf cannot
SIGSTOPa server it does not own — a systemd one belongs toclickhouse— and does not need to: an already-running process has no startup for the pause to protect. The profile starts when the collectors open. Profiling still needsperfwithCAP_PERFMON, which its file capabilities usually carry;vperf doctorchecks it and prints thesetcapline for your host if not.
Expect a few seconds of post-processing per query on top of the query itself:
perf.data for a one-second window over the whole server is ~100 MB, and both
the flush and the perf script pass read all of it. It is deleted afterwards, so
a sweep does not leave it behind; --keep-perf-data keeps it.