Skip to content

feat: sleep when blocked, SPSC bulk transfer, MPMC pow2 indexing - #10

Merged
jadecubes merged 4 commits into
mainfrom
feat/industrial-gap-features
Aug 25, 2026
Merged

jadecubes merged 4 commits into
mainfrom
feat/industrial-gap-features

Conversation

@jadecubes

@jadecubes jadecubes commented Aug 24, 2026 •

Copy link
Copy Markdown
Owner

Summary

Three features closing the gap to the industrial queues, with the two designs that measured worse documented as findings:

  • Blocked ops sleep instead of spinning. Blocking push/pop yield briefly, then sleep with doubling backoff (new cq/backoff.hpp): 2 s blocked = 10 ms of CPU, down from 1975 ms, wakeup bounded at 1 ms. Rejected on measurement: a futex gate (C++20 atomic wait/notify eventcount) cost 70% of SPSC pair throughput — the publisher-side waiter check after a release store drains the store buffer; and inlining the sleep machinery into push/pop cost 4× on the round trip by blowing the inliner budget — hence the noinline cold Backoff::wait(). Honest residual: continuous 1p+1c rendezvous runs ~9% below pure spin (547M vs 602M ops/s); round trip unchanged at 1.80G ops/s.
  • SPSC bulk transfer. try_push_n/try_pop_n: up to a std::span's worth of items per call, one index publish per batch — the lever the results' cached-index section named. 1.69G ops/s at batch 64, 3.1× the single-op pair.
  • MPMC pow2 indexing. ticket→slot is one AND when capacity is a power of two (division otherwise; capacity stays arbitrary). Within noise on M2; kept as strictly cheaper.

Files changed & why

File Change Why
include/cq/backoff.hpp new The wait policy: yield × 64, then doubling sleep 4 µs → 1 ms; noinline wait() keeps callers inlinable
include/cq/spsc_queue.hpp/.ipp modified Blocking loops use Backoff; try_push_n/try_pop_n (span API, single publish per batch); docs
include/cq/mpmc_queue.hpp/.ipp modified Blocking loops use Backoff; slot_index() with pow2 mask; docs
tests/spsc_queue_test.cpp modified 6 bulk tests (partial fill/drain, wrap, close, move-only), written first and watched failing
bench/queue_bench.cpp modified SpscQueue/bulk64_throughput benchmark
README.md modified Blocked-thread row now says the rings spin briefly then sleep; bulk-ops row; Performance bullet for the ~9% residual and the bulk figure
docs/results.md modified New v3.1 section: full writeup incl. both rejected designs and the cross-session round-trip control note

Post-review updates (after the #11 restructure landed on main)

  • Merged main in and resolved the README conflict: the results writeup now lives in docs/results.md (v3.1 section + milestone-map row); the restructured README got the integration edits listed above instead of a results section.
  • Review round: bulk ops take std::span<T> instead of pointer + length (A/B benched: span 1.64G ± 0.03 vs pointer 1.60G ± 0.04 — no cost); Backoff::wait() moved out of the class body per STYLE.md and the header gained the required /// docs; slot_index's comment corrected for capacity 1 (pow2 with mask 0 — code was already right); unused <algorithm>/<chrono>/<thread> includes dropped from the .ipps.

Test plan

  • 65/65 under ThreadSanitizer (existing blocked/close-wake contract tests now exercise the sleep paths); 66/66 in Release incl. bench smoke — re-run after the merge and the span change
  • clang-format clean; clang-tidy 0 findings
  • Sleep A/B probe (2 s blocked consumer): main 1974.7 ms CPU → this branch 10.1 ms
  • Hot-path parity verified by standalone round-trip A/B: main 1.36 ns, this branch 1.32 ns
  • Bulk figure reproduced after the span change: 1.64G ± 0.03 ops/s, within 1σ of the published 1.69G ± 0.09G

🤖 Generated with Claude Code

Debra and others added 4 commits August 25, 2026 07:06
Blocking push/pop yield briefly then sleep with doubling backoff
(cq/backoff.hpp): 2s blocked drops from 1975ms to 10ms of CPU, wakeup
bounded at 1ms. Two designs were measured and rejected on the way — a
futex gate cost 70% of SPSC pair throughput (publisher-side check
drains the store buffer), and inlining the sleep machinery into
push/pop cost 4x on the round trip (inliner budget) — hence the
noinline cold Backoff::wait(). Remaining cost: ~9% on the continuous
1p+1c rendezvous; round trip unchanged at 1.80G ops/s.

try_push_n/try_pop_n move a batch with one index publish: 1.69G ops/s
at batch 64, 3.1x the single-op pair.

MpmcQueue maps ticket->slot with an AND when capacity is a power of
two; within noise on M2, kept as strictly cheaper.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ve-only test

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ructure

PR #11 moved per-milestone results to docs/results.md and rewrote the
README around the three queues, so this branch's README section conflicts.
Resolution:

- The new results writeup moves to docs/results.md as a v3.1 section, with
  a row in the milestone map and an explicit note on the round-trip
  control level differing across sessions.
- README (main's version) is updated where this branch made it stale: the
  "A waiting thread" row now says the rings spin briefly then sleep, the
  three-queues table gains a bulk-ops row, and Performance gains a bullet
  for the ~9% sleeping residual and the 1.69G bulk figure.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0133qvk9V3aREYsxvMCbZ9Am
…hygiene

- try_push_n/try_pop_n take std::span<T> instead of pointer + length: the
  span carries both, call sites shrink, and the A/B benchmark shows no
  cost (span 1.64G ± 0.03 vs pointer 1.60G ± 0.04 ops/s at batch 64).
- Backoff::wait() moves out of the class body per STYLE.md ("no function
  bodies in class definitions") and gains the /// contract comment
  include/cq/ requires; the constants get /// one-liners.
- slot_index's comment claimed the mask is set iff capacity is a power of
  two — false at capacity 1 (2^0, mask 0, modulo path). Comment now says
  what happens there; the code was already right.
- Drop includes the switch from yield loops to Backoff left unused:
  <chrono>/<thread> in spsc_queue.ipp, <algorithm>/<chrono>/<thread> in
  mpmc_queue.ipp.

Verified: 65/65 under TSan, 66/66 in Release; clang-format clean;
clang-tidy 0 findings; bulk benchmark within 1 sigma of the published
1.69G ± 0.09G.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0133qvk9V3aREYsxvMCbZ9Am
@jadecubes jadecubes self-assigned this Aug 25, 2026
@jadecubes
jadecubes merged commit ca26527 into main Aug 25, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant