Repository navigation
3.x.x: stop SharedLog losing events whose callsite another test reached first, and order the block-only poll faults after the enqueue (#649) - #653
Merged
Conversation
…ed first, and order the block-only poll faults after the enqueue (#649) commit_reconcile_block_only_poll_errors_are_retried_then_unknown failed under the full parallel --lib run because its log lost the "block-only disposition poll failed; retrying" warning, not because of the wall clock. Every recorded failure (three from 2026-10-01, and one reproduced here on eedf331) shows the first poll failing at exactly the 5 s lock_timeout and the final warning carrying phase="poll-error", so the polls did fail and retry in time; only that one event was missing. tracing caches one interest per callsite for the process. While a single dispatcher is registered, a callsite's first hit asks only the hitting thread's default. miner_submit::acquire_tests run block-only polls against a minimal schema, so their polls fail and reach the same callsite on threads without a subscriber. When that happened while this test's SharedLog was the only registered dispatcher, the callsite was cached `never` and the SharedLog dropped the event. A diagnostic run showed the acquire test's first hit on a thread with the global (none) default, between this test's submit and its first failed poll. SharedLog::dispatch now first installs, once, a global default that records nothing and answers every callsite `sometimes`. There are then always two registered dispatchers, and it is the default of any thread without its own, so no callsite can be cached `never` and each event asks the current dispatcher. A unit test reaches a fresh callsite on a thread without a subscriber while one SharedLog is live: it fails without the default and passes with it. The test also no longer takes its table lock after watching the candidate become pending, which left a 3 s wall-clock window (bound minus lock_timeout) for the first poll to fail. A wrapper on the probe floor that every poll calls fails each poll with SQLSTATE 55P03 (lock_not_available) once the candidate is committed, and counts the failures in a sequence; the duplicate probe before the enqueue passes through. The test asserts at least two failed polls from the database as well as the logged retry, and still that the answer is ledger-outcome-unknown at the bound, never a fabricated outcome, and that the later confirmation credits the proof once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Member
Author
|
Re-running CI on the current |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #649.
commit_reconcile_block_only_poll_errors_are_retried_then_unknownwas not losing a wall-clock race: its log lost one warning. Another lib test reached the sametracingcallsite first, on a thread without a subscriber, andtracingcached the callsite asnever. This PR makesSharedLogimmune to that, and also orders the test's injected poll failures after the enqueue in the database instead of by wall clock. The change is test-only. No timeout, retry or production code changes, and every original assertion is kept.Problem
Every recorded failure looks the same: three from 2026-10-01 (PR #643's run, PR #645's worker, #570's work) and one reproduced here at
eedf331b. The first disposition poll fails at exactly the 5 slock_timeout(sqlx logselapsed=5.000s). The final warning carriesphase="poll-error". So the polls did fail and retry well inside the 8 s bound. Only theblock-only disposition poll failed; retryingwarning is missing from the captured log, so thetext.contains(...)assertion fails.The cause is in
tracing-core0.1.36:tracingcaches one interest per callsite for the whole process.Rebuilder::JustOne→get_default). A thread without a subscriber therefore cachesnever.coordinator::miner_submit::acquire_testsrun block-only polls against a minimal schema (nooffer_outcomecolumn), so their polls fail and reach the same callsite. They do not take the coordinatorTEST_LOCK, and they have no subscriber.SharedLogis the only registered dispatcher, the callsite is cachednever, and theSharedLogsilently drops this test's warning.A temporary diagnostic run at the warning site showed the order: this test submits, then an acquire test's
checkout:…share hits the callsite on a thread whose default is the global none (column "offer_outcome" does not exist), then this test's poll fails. At that point the global max level wasWARNand the thread's default was this test's scoped dispatcher, so the drop came from the cached interest. That is why the test passes alone and fails only in the full parallel run.The issue's suspected window was real but was not what failed. The old test took
LOCK TABLEonly after it saw the candidate pending, so the first poll had to fail bybound − lock_timeout= 3 s after the submit. It never exceeded that in the recorded failures.Change
miner_tests/mod.rs:SharedLog::dispatch()first installs a global default subscriber,Undecided, once.Undecidedrecords nothing and answers every callsitesometimes, with anOFFlevel hint so it never raises the global maximum.never, and each event asks the current dispatcher.mainsets a global default, so no lib test conflicts with it.miner_tests/mod.rs: new testshared_log_captures_a_callsite_first_reached_without_a_subscriber. It reaches a fresh callsite on a thread without a subscriber while oneSharedLogis live, then again under thatSharedLog.commit_reconcile_tests.rs: the test no longer takes its table lock after polling for the pending candidate.qbit_prism_share_probe_floor(), which every disposition poll calls, in the same rename-and-replace way astests/b574_ack_cap.rs.55P03(lock_not_available, the class a lock timeout raises) and counts itself in a sequence. The duplicate probe before the enqueue finds no candidate and passes through.ledger-outcome-unknown(never a fabricated outcome), answered no earlier than the 8 s bound, and credited exactly once by the later confirmation, after the wrapper is removed.Tests
On a private PostgreSQL 16 (
PRISM_TEST_DATABASE_URLset,PRISM_TEST_REQUIRE_INTEGRATION=1for the single-test runs), one test binary at a time, no synthetic load.Reproduction.
eedf331bcargo test --locked -p qbit-prism-server --lib, full parallelthe failed polls were not retried and logged, withphase="poll-error"present and the retry warning absentUndecided::install()commented out--exact coordinator::miner_tests::shared_log_captures_a_callsite_first_reached_without_a_subscriberthe SharedLog lost an event whose callsite was first reached elsewherepg_sleepbefore the enqueue (beyond the old 3 s window)--exact …poll_errors_are_retried_then_unknownSuites (this PR):
cargo test --locked -p qbit-prism-server --lib, full parallel, 3 runs back to back: 623 passed, 0 failed, 2 ignored each time (113.4 s, 126.1 s, 118.6 s). Host load average about 3.5–4.5.--exact coordinator::commit_reconcile_tests::commit_reconcile_block_only_poll_errors_are_retried_then_unknown: 2 of 2 passed, 8.5 s each.cargo fmt --all -- --check: clean.cargo clippy --locked -p qbit-prism-server --all-targets -- -D warnings: clean.python3 scripts/check_prism_pool_acquires.pyandpython3 scripts/check_gate_env_reads.py: pass.The PostgreSQL test is already in
test/prism-gated-tests.txt. The new unit test needs no database, so the list andtest/e2e-scenarios.tomlare unchanged.Not run here and why
The build host has no qbitd and no Docker. Nothing in this change needs either; the lib tests above are the whole affected surface.
Merge notes
3.x.xateedf331b. Test-only, no ordering constraints.d2_below_target_tests.rshas its ownSharedLog, which does not install the default. It is protected only when aminer_testsSharedLoghas already installed it in the same binary. It can switch to the shared helper in a follow-up.deadlock detected. It did not occur in the passing runs.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.