Skip to content

bench: sweep the non-blocking path across capacity, on one clock - #14

Merged
jadecubes merged 6 commits into
mainfrom
fix/benchmark-success-accounting
Aug 27, 2026
Merged

jadecubes merged 6 commits into
mainfrom
fix/benchmark-success-accounting

Conversation

@jadecubes

@jadecubes jadecubes commented Aug 27, 2026 •

Copy link
Copy Markdown
Owner

What this adds

main only benchmarks the blocking push/pop path. This adds the non-blocking one: BM_QueueTryThroughput<Queue>, swept across capacities 1, 2, 8, 64 and 1024 for MutexQueue, SpscQueue and MpmcQueue.

Failed attempts are retried inside the timed iteration, so they cost time but never count as transferred work — items/s stays the same unit as the blocking rows. The pressure that caused them is reported separately as retries/op, split into push_retries/push and pop_retries/pop.

Result

5 invocations × 10 repetitions, load 3.5–6.1, items/s as mean (min–max):

capacity MutexQueue SpscQueue MpmcQueue Spsc/Mutex
1 5.6 (5.3–6.0) 7.8 (7.6–7.9) 6.9 (6.7–7.2) 1.3–1.5×
2 9.6 (9.2–10.0) 15.1 (14.5–15.5) 14.1 (13.6–14.4) 1.5–1.6×
8 27.2 (25.8–28.9) 61.3 (59.9–62.2) 50.7 (48.6–54.1) 2.1–2.4×
64 20.1 (19.5–20.5) 521.1 (501.1–531.1) 186.4 (183.9–190.4) 24.6–26.9×
1024 54.3 (53.8–55.2) 622.6 (616.3–627.4) 195.2 (189.9–198.3) 11.3–11.7×

The lock-free advantage is not a constant. It is small while the ring is shallow and opens up once a producer can run ahead. retries/op tracks it: 1.05–1.30 at capacity 1, 0.0005–0.0054 at 1024.

Read the table with the retry policy in mind. The loop yields after each failure, which only helps MutexQueue — its try_ calls take a blocking lock_guard, so an unyielding spinner barges the lock back from the counterpart that would have made room. The lock-free queues have no lock to barge, so the yield is pure overhead there. Removing it at capacity 1 costs MutexQueue ~10× and gains each lock-free queue 2–3×, moving Spsc/Mutex from under 2× to around 40×.

The yield stays (without it the mutex row degenerates, and a 2-vCPU runner would be worse than this 12-core box), but it is now named in the helper and in the sweep's comment. Ratios rather than rates throughout: absolute figures move 10–25% with machine load.

Two defects fixed along the way

Clock. The five single_thread_roundtrip registrations set no UseRealTime(), so they reported items/s per CPU-second while every threaded row is per wall-second — despite a comment claiming the units matched. Small (1.004–1.015) but two different quantities.

Lint coverage. HeaderFilterRegex covered only include/cq/, so tests/queue_test_util.hpp and the new bench/try_operation.hpp got no clang-tidy diagnostics at all. Widening it exposed a second problem: the filter matches a path segment above the repo root, so a checkout under /home/runner/work/bench/bench/ lints benchmark.h under WarningsAsErrors (reproduced at 187 errors). Fixed by declaring benchmark SYSTEM, as concurrentqueue and tbb already are.

One open question

MutexQueue is slower at capacity 64 than at capacity 8. Paired sampling — both points inside the same invocation, eight invocations — gives 8/8 in the same direction, difference +9.47 ± 2.36 M/s. It is real, and I have no explanation for it.

Scope

Only the cq queues are swept: neither external adapter has a try_pop, and moodycamel is unbounded so it has no capacity to vary.

Files

File +/− What changed
bench/queue_bench.cpp +138 −10 Adds BM_QueueTryThroughput<Queue> — same thread split as the blocking benchmark, but each iteration retries until it succeeds and accumulates the failures, then publishes three kAvgIterations counters. Adds setup_queue_at_capacity<Queue>, which reads the capacity from the registered Arg so one benchmark can sweep it, and hoists the prefill both setups share into install_half_full_queue<Queue>. Registers the sweep three times against a named capacity_sweep list. Adds ->UseRealTime() to the five single_thread_roundtrip registrations.
bench/try_operation.hpp +51 New. count_failures_until_success(op) — retry until success, return the failure count, std::this_thread::yield() after each failure. Constrained with invocable<Operation&> && convertible_to<invoke_result_t<Operation&>, bool> rather than std::predicate, which would demand equality-preserving invocation this deliberately violates. Most of the file is the rationale for the yield: what it buys, what it costs, and why the ratios and not the rates are quoted.
bench/CMakeLists.txt +16 Adds the bench_retry_pressure ctest, which runs the capacity-1 sweep point and asserts retries/op is non-zero on the _mean row. Separate from bench_smoke because PASS_REGULAR_EXPRESSION suppresses CTest's exit-code check, and bench_smoke should keep it.
tests/try_operation_test.cpp +23 New. One test: an operation that succeeds on its third call is invoked exactly three times and reports exactly two failures. Catches a helper that gives up early — which the ctest gate cannot, since such a helper still reports non-zero.
tests/CMakeLists.txt +4 −1 Compiles the new test and points the test target at bench/ so it can include the helper.
.clang-tidy +7 −1 HeaderFilterRegex: 'include/cq/.*' → '(^|/)(include/cq|bench|tests)/'. Both tests/queue_test_util.hpp and the new bench/try_operation.hpp sit outside include/cq/ and were receiving no diagnostics at all.
CMakeLists.txt +8 −1 Adds SYSTEM to benchmark's FetchContent_Declare. Required by the row above: it is the only dependency arriving as a plain -I, so it is the only one the widened filter can reach.

Verification

Release ctest 68/68 (66 before), TSan 66/66, clang-format and clang-tidy clean. Stubbing the helper to return 0; fails bench_retry_pressure while bench_smoke still passes.


Originally opened against feat/mutex-queue, a branch already merged as #2 that main has since moved past. Rebuilt on main and re-targeted; pre-rebuild head was da2874c.

🤖 Generated with Claude Code

@jadecubes jadecubes self-assigned this Aug 27, 2026
@jadecubes
jadecubes force-pushed the fix/benchmark-success-accounting branch from e72a97d to da2874c Compare August 27, 2026 09:08
jadecubes pushed a commit that referenced this pull request Aug 27, 2026
Decision 1 settled: SpscQueue ships try_push/try_pop only, with no blocking
push()/pop() and no close(). Blocking would need a condition variable, which
needs a mutex, which contaminates the claim v2 exists to test. The fairness
half is already handled on the v1 side -- PR #14 added a MutexQueue try_
throughput sweep -- so the comparison is like-for-like without giving v2 a
lock on its slow path.

Folded in two findings from that PR that constrain v2's benchmark: a retry
loop without backoff makes shallow capacities swing ~400x between runs, and
every registration must set UseRealTime() or its throughput is normalised by
a different clock than the rows beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a try_ counterpart to the blocking throughput harness and sweeps it
over capacity for the three cq queues, plus the retry diagnostics needed to
read the result.

Failed attempts are retried inside the timed iteration, so they cost time
but never count as transferred work: items/s stays the same unit as the
blocking rows. What produced those failures is reported separately as
retries/op, with push_retries/push and pop_retries/pop splitting it by side
-- the aggregate alone cannot say which side is under pressure, which is the
thing a capacity sweep exists to expose.

The retry loop yields after each failure. That is load-bearing, not
politeness: without it a spinning side keeps barging the lock back from its
counterpart, which then cannot make the progress that would let the spinner
succeed. At capacity 1 or 2 every op depends on the counterpart, so one
starvation episode dominates a whole run -- measured on an earlier draft,
the reported rate at capacity 1 swung 396x across five identical
invocations.

What the sweep shows (10 repetitions, load ~4.2, items/s):

    capacity   MutexQueue   SpscQueue   MpmcQueue   Spsc/Mutex
           1        5.5M        7.8M        6.7M        1.42x
           2        8.8M       14.4M       13.5M        1.64x
           8       24.6M       59.8M       47.9M        2.43x
          64       19.6M      469.6M      179.6M       23.95x
        1024       53.9M      580.9M      180.2M       10.78x

The lock-free advantage is not a constant: it nearly vanishes at capacity 1
and only opens up once the ring is deep enough for a producer to run ahead.
A shallow queue makes every op wait on its counterpart no matter how the
waiting is implemented, so both designs converge on the cost of a cross-core
handoff. retries/op tracks it exactly -- 1.07-1.31 at capacity 1, 0.001-0.005
at 1024. Quoting a single lock-free speedup without naming the depth it was
measured at says very little.

Also fixes a clock inconsistency this sweep would otherwise inherit: the
five single_thread_roundtrip registrations set no UseRealTime(), so Google
Benchmark normalised their rate counter by CPU time while every threaded row
used wall time. The comment on their SetItemsProcessed claims the unit is
kept "identical to the threaded benchmark so the rates compare directly" --
true of the item count, false of the clock.

The sweep covers only the cq queues: moodycamel is unbounded and so has no
capacity to vary, and the tbb adapter has no try_pop yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jadecubes
jadecubes force-pushed the fix/benchmark-success-accounting branch from da2874c to 69d58dd Compare August 27, 2026 09:59
@jadecubes
jadecubes changed the base branch from feat/mutex-queue to main August 27, 2026 10:02
@jadecubes jadecubes changed the title bench: count only successful try operations bench: sweep the non-blocking path across capacity, on one clock Aug 27, 2026
@jadecubes jadecubes closed this Aug 27, 2026
@jadecubes jadecubes reopened this Aug 27, 2026
Debra and others added 5 commits August 27, 2026 19:00
…the queues

Review follow-ups on the sweep this branch added. The headline reading was
wrong, and the code comments oversold one queue's behaviour as all three.

The yield in count_failures_until_success was justified as stopping a
spinner from barging the lock back from its counterpart. That mechanism
exists only for MutexQueue, which takes a blocking lock_guard inside
try_push/try_pop. SpscQueue and MpmcQueue have no lock to barge, so for
them the yield is pure overhead. Measured at capacity 1:

    MutexQueue   0.58 M/s without -> 5.11 M/s with   (yield buys 8.8x)
    SpscQueue   20.9  M/s without -> 7.84 M/s with   (yield costs 2.7x)
    MpmcQueue   21.1  M/s without -> 6.84 M/s with   (yield costs 3.1x)

So the conclusion this sweep appeared to support -- that the lock-free
advantage nearly vanishes at shallow depth -- is a property of the retry
policy, not of the queues. Spsc/Mutex at capacity 1 is ~1.5x with the yield
and ~36x without it. The yield stays, because without it the mutex row
degenerates to 155 retries per op and a two-vCPU CI runner would be far
worse than this 12-core box; but it is now named where a reader will meet
it, and the sweep's own comment says any ratio taken from these rows is a
statement about the policy too.

"Every value is a power of two because MpmcQueue masks its indices" was
false: mpmc_queue.ipp falls back to % when capacity is not a power of two,
and QueueContract/Mpmc.FillsToExactlyCapacity exercises capacity 3. The
masking is a fast path, not a constraint, and the sweep values are a choice.

Cleanups alongside: hoist the two prefill loops, which had drifted into two
different failure idioms fifteen lines apart, into install_half_full_queue;
name the sweep list once instead of repeating it at three registrations,
which also retires the NOLINT(readability-magic-numbers) bracket, since
clang-tidy exempts literals in a const initializer; use Google Benchmark's
documented counters["name"] = Counter(...) idiom so the counter names land
at the left margin; trim two rationale paragraphs that restated what the
code beside them already showed; and note that an odd producer/consumer
split would misreport the per-side counters but deadlocks the harness first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The measured table this comment carried was itself the mistake it was
added to fix. Re-measuring at a different machine load moved every one of
its six absolute figures by 10-25%: MutexQueue 0.58 -> 0.46 M/s without the
yield, SpscQueue 20.9 -> 18.2, MpmcQueue 21.1 -> 16.6, and the headline
Spsc/Mutex ratio from ~36x to ~40x. The "155 retries per op" measured 194.

The directions and rough magnitudes reproduce on every run; the rates do
not. So the comment now states only what survives — removing the yield
costs MutexQueue about a factor of ten and gains each lock-free queue two
to three — and says why no rates are quoted.

Same correction in the sweep's comment: Spsc/Mutex at capacity 1 is under
2x with the yield and around 40x without, rather than the ~1.5x / ~36x
taken from a single run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ew header

Round-2 review follow-ups.

"Powers of two keep MpmcQueue on its mask fast path" is false at capacity 1
— the sweep's shallowest point, and the one bench_retry_pressure keys on.
has_single_bit(1) is true, so the ctor sets mask_ = 1 - 1 = 0, and
slot_index() gates on mask_ != 0, sending capacity 1 down the % path: a
runtime modulo by a runtime divisor, not the fast path the comment
promises.

"An odd split would misreport, but it deadlocks this harness first" was
overstated twice. It is a spin-with-yield livelock, not a deadlock. And
"first" is conditional: with N iterations per thread against a prefill of
capacity/2, an odd split completes normally whenever the surplus pops fit
inside the prefill — bench_smoke's own regime (1x at capacity 1024, prefill
512) is exactly that, so there it would report silently wrong counters
rather than hang.

bench/try_operation.hpp is the first header this repo has placed outside
include/cq, and so the first one clang-tidy never sees: HeaderFilterRegex
was scoped to 'include/cq/.*' and CI lints only *.cpp, so a diagnostic in
the new header was silently dropped. Widened to '(include/cq|bench)/.*'.
Verified both directions: an injected BadlyNamedVar in that header now
produces readability-identifier-naming, and with the probe removed every
tracked .cpp is still clean under the wider filter.

Also record why capacity_sweep is const rather than only why it is not
constexpr — the const is what makes clang-tidy exempt its literals, so
removing it as redundant would resurrect the magic-numbers error the same
commit deleted a NOLINT for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round-3 review follow-ups. Both are claims the previous commit introduced
while fixing round 2 — the third round running in which a fix has produced
a new false statement about the same subject.

"The first header this repo has placed outside include/cq" was wrong three
times over: in the .clang-tidy comment, the commit message, and the PR body.
tests/queue_test_util.hpp already sat outside it on main, and the widened
regex left it out — so the gap the commit claimed to close was only half
closed, and the stated rationale ("without it that header receives no
diagnostics at all") applied verbatim to a header the fix skipped. The
filter now covers include/cq, bench and tests, and is anchored: it was an
unanchored substring match, and googletest and benchmark are not declared
SYSTEM here, so a checkout under a path containing a bench/ segment would
have started linting dependency headers. Verified an injected bad name in
both headers now produces diagnostics, and that every tracked .cpp is still
clean with the probes removed.

"The divisor would need to be the thread count, not 2" was also wrong.
kAvgIterations divides by the iteration total across all threads, so the
scale is threads/producers on the push side and threads/consumers on the
pop side; the thread count is right only when a side has exactly one
thread. Measured at threads:3, capacity:1024: pop_retries/pop reported
1.30208m against a ground truth of 976.6u, over by exactly 4/3.

The livelock qualifier was also lost between the commit message and the
comment. The discriminator is iterations against prefill, not depth: under
the registered MinTime, iterations-per-thread grows far past any prefill,
so an odd split livelocks at every capacity including the deepest. Only an
iteration-capped smoke run completes, and only there would it report
silently wrong numbers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…PR opened

The anchoring rationale in the previous commit was false, and the thing it
claimed to prevent is a real breakage this PR introduced.

`(^|/)` only excludes segment-substring matches like microbench/. It still
matches a bench/ segment anywhere above the repo root — so widening
HeaderFilterRegex to bench|tests exposed dependency headers to
WarningsAsErrors wherever such a segment appears in the path. A GitHub
Actions repo named `bench` checks out to /home/runner/work/bench/bench/,
and `cmake -B tests/build` does it locally. Reproduced by relocating the
compile database under /tmp/bench/link: 199 hard errors out of
benchmark/benchmark.h. main's include/cq/.* filter was immune; this branch
was not.

The fix belongs in CMakeLists, not the regex. benchmark is the only
dependency arriving as a plain -I: concurrentqueue and tbb already declare
SYSTEM, and googletest self-marks its interface SYSTEM regardless. Adding
SYSTEM to its FetchContent_Declare takes the same relocated build from 199
diagnostics to 0, while the ordinary path stays clean and every tracked
.cpp still lints without a diagnostic. The .clang-tidy comment now says
dependencies stay out because they are SYSTEM, which is true, instead of
crediting the anchor, which is not.

Also replace the odd-split sentence that three rounds have now rewritten
with the single condition it was circling: an odd split completes only
while (consumers - producers) * iterations-per-thread <= capacity / 2. The
"every capacity under MinTime" case and the smoke-run case both fall out of
it, and the previous phrasing was wrong at capacity 1, where the prefill is
0 and an odd split hangs even under an iteration cap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jadecubes
jadecubes merged commit ffac12e into main Aug 27, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant