Split out of #39 by the 2026-09-26 0330 run. #39 fixed the within-run noise (min-of-5 sampling + a noiseRatio withhold). This is the residue it could not fix, measured on the same day, same machine.
Finding
After #39 landed (d888742), the gate's spurious findings dropped sharply and became diagnosable — but they did not stop:
Measured openbook_decode across 8 clean-tree runs post-fix: 6.05, 6.10, 6.25, 6.40, 6.53, 6.73, 6.12, 5.54 ms — against an anchor of 5.28–5.73 ms and the section's 10% default allowance. So it trips by 0–15% while the arm's own within-run spread stays tight (1.01–1.38x, i.e. under the 1.5x withhold threshold).
The instrument is not lying about the run. It is comparing against a stale bar: perf/results-accepted.json was last anchored on 2026-09-19 and this machine is simply running this arm ~8–15% slower today. Nothing regressed; the reference point aged.
Why noiseRatio structurally cannot catch this
noiseRatio measures dispersion inside one run. This failure is a shift in the level between runs — the anchor was recorded on a different day, at a different thermal/clock state, possibly with a different working set. Tight samples are entirely consistent with a level that has drifted, so the withhold correctly stays silent and the comparison correctly fires. The gate is doing its job; its reference is wrong.
This is also why the anchor ratchets: each PASS rewrites the anchor, and a lucky-fast PASS records a bar stricter than the machine can hold, which manufactures a failure on the very next run. Observed directly in artifacts/automation-receipts/2026-09-26/anchor-drift-from-noise.txt — one lucky pre-fix PASS moved every anchor value 2–19% (cli_startup_help 25.2 → 30.0, i.e. the bar got worse).
Decision question (evaluate — implementation is NOT the todo)
Pick the anchor-stability policy, with evidence:
- drift-aware tolerance — compare against
anchor.currentMs * (1 + max(allow, anchor.spread - 1)), i.e. never hold a bar tighter than the anchor's own recorded uncertainty;
- anchor re-baselining cadence — a periodic (weekly / Node-version / machine-change) re-seed, with the re-seed itself recorded so it is never silent;
- median-of-K anchors — keep the last K accepted runs per section and compare against their median, which is far less hostage to one lucky or one unlucky day;
- accept the trip — document that the anchor is deliberately strict and a recurring single-section trip on
openbook_decode is expected operator noise.
The deciding constraint: the run-over-run anchor exists to catch "a modest since-last-week regression" (#29). Any of these must keep catching a real 20%+ regression on a stable machine, so the fix cannot simply be a wider allowance — that is the band-aid #39 removed from the sampling side and would reintroduce it on the comparison side.
Evidence
Note: the per-run artifacts for the intermediate REPS=3 iteration of #39 were overwritten by the REPS=5 sweep; their numbers survive only in the run receipt's log (artifacts/automation-receipts/2026-09-26/missing-runid-attempt-1-libby-archiver-0330.md).
Split out of #39 by the 2026-09-26 0330 run. #39 fixed the within-run noise (min-of-5 sampling + a
noiseRatiowithhold). This is the residue it could not fix, measured on the same day, same machine.Finding
After #39 landed (
d888742), the gate's spurious findings dropped sharply and became diagnosable — but they did not stop:openbook_decode(localized).Measured
openbook_decodeacross 8 clean-tree runs post-fix: 6.05, 6.10, 6.25, 6.40, 6.53, 6.73, 6.12, 5.54 ms — against an anchor of 5.28–5.73 ms and the section's 10% default allowance. So it trips by 0–15% while the arm's own within-run spread stays tight (1.01–1.38x, i.e. under the 1.5x withhold threshold).The instrument is not lying about the run. It is comparing against a stale bar:
perf/results-accepted.jsonwas last anchored on 2026-09-19 and this machine is simply running this arm ~8–15% slower today. Nothing regressed; the reference point aged.Why
noiseRatiostructurally cannot catch thisnoiseRatiomeasures dispersion inside one run. This failure is a shift in the level between runs — the anchor was recorded on a different day, at a different thermal/clock state, possibly with a different working set. Tight samples are entirely consistent with a level that has drifted, so the withhold correctly stays silent and the comparison correctly fires. The gate is doing its job; its reference is wrong.This is also why the anchor ratchets: each PASS rewrites the anchor, and a lucky-fast PASS records a bar stricter than the machine can hold, which manufactures a failure on the very next run. Observed directly in
artifacts/automation-receipts/2026-09-26/anchor-drift-from-noise.txt— one lucky pre-fix PASS moved every anchor value 2–19% (cli_startup_help25.2 → 30.0, i.e. the bar got worse).Decision question (evaluate — implementation is NOT the todo)
Pick the anchor-stability policy, with evidence:
anchor.currentMs * (1 + max(allow, anchor.spread - 1)), i.e. never hold a bar tighter than the anchor's own recorded uncertainty;openbook_decodeis expected operator noise.The deciding constraint: the run-over-run anchor exists to catch "a modest since-last-week regression" (#29). Any of these must keep catching a real 20%+ regression on a stable machine, so the fix cannot simply be a wider allowance — that is the band-aid #39 removed from the sampling side and would reintroduce it on the comparison side.
Evidence
artifacts/automation-receipts/2026-09-26/noise/results-fix5reps-{1..8}.json— post-perf: the bench gate false-trips on a clean tree in 67% of runs — --quick arms time most sections with a single sample, so run-to-run spread (up to 3.3x) dwarfs the 10-35% allowances #39 clean-tree runsartifacts/automation-receipts/2026-09-26/noise/results-noise{1..6}.json— pre-perf: the bench gate false-trips on a clean tree in 67% of runs — --quick arms time most sections with a single sample, so run-to-run spread (up to 3.3x) dwarfs the 10-35% allowances #39 clean-tree runsartifacts/automation-receipts/2026-09-26/anchor-drift-from-noise.txt— the ratchet, caughtd888742(the fix this splits out of)Note: the per-run artifacts for the intermediate
REPS=3iteration of #39 were overwritten by theREPS=5sweep; their numbers survive only in the run receipt's log (artifacts/automation-receipts/2026-09-26/missing-runid-attempt-1-libby-archiver-0330.md).