Skip to content

perf: run-over-run anchor has no drift policy — a 6-day-old bar made openbook_decode trip 6 of 8 clean-tree runs, and one lucky PASS ratchets the bar stricter #40

Description

@JavaGT

Split out of #39 by the 2026-09-26 0330 run. #39 fixed the within-run noise (min-of-5 sampling + a noiseRatio withhold). This is the residue it could not fix, measured on the same day, same machine.

Finding

After #39 landed (d888742), the gate's spurious findings dropped sharply and became diagnosable — but they did not stop:

Measured openbook_decode across 8 clean-tree runs post-fix: 6.05, 6.10, 6.25, 6.40, 6.53, 6.73, 6.12, 5.54 ms — against an anchor of 5.28–5.73 ms and the section's 10% default allowance. So it trips by 0–15% while the arm's own within-run spread stays tight (1.01–1.38x, i.e. under the 1.5x withhold threshold).

The instrument is not lying about the run. It is comparing against a stale bar: perf/results-accepted.json was last anchored on 2026-09-19 and this machine is simply running this arm ~8–15% slower today. Nothing regressed; the reference point aged.

Why noiseRatio structurally cannot catch this

noiseRatio measures dispersion inside one run. This failure is a shift in the level between runs — the anchor was recorded on a different day, at a different thermal/clock state, possibly with a different working set. Tight samples are entirely consistent with a level that has drifted, so the withhold correctly stays silent and the comparison correctly fires. The gate is doing its job; its reference is wrong.

This is also why the anchor ratchets: each PASS rewrites the anchor, and a lucky-fast PASS records a bar stricter than the machine can hold, which manufactures a failure on the very next run. Observed directly in artifacts/automation-receipts/2026-09-26/anchor-drift-from-noise.txt — one lucky pre-fix PASS moved every anchor value 2–19% (cli_startup_help 25.2 → 30.0, i.e. the bar got worse).

Decision question (evaluate — implementation is NOT the todo)

Pick the anchor-stability policy, with evidence:

  • drift-aware tolerance — compare against anchor.currentMs * (1 + max(allow, anchor.spread - 1)), i.e. never hold a bar tighter than the anchor's own recorded uncertainty;
  • anchor re-baselining cadence — a periodic (weekly / Node-version / machine-change) re-seed, with the re-seed itself recorded so it is never silent;
  • median-of-K anchors — keep the last K accepted runs per section and compare against their median, which is far less hostage to one lucky or one unlucky day;
  • accept the trip — document that the anchor is deliberately strict and a recurring single-section trip on openbook_decode is expected operator noise.

The deciding constraint: the run-over-run anchor exists to catch "a modest since-last-week regression" (#29). Any of these must keep catching a real 20%+ regression on a stable machine, so the fix cannot simply be a wider allowance — that is the band-aid #39 removed from the sampling side and would reintroduce it on the comparison side.

Evidence

Note: the per-run artifacts for the intermediate REPS=3 iteration of #39 were overwritten by the REPS=5 sweep; their numbers survive only in the run receipt's log (artifacts/automation-receipts/2026-09-26/missing-runid-attempt-1-libby-archiver-0330.md).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions