feat(server): score scraper lanes on frozen benchmarks and store the trend - #3535
Merged
Merged
Conversation
…he trend Nothing measured a lane against a fixed input, so a lane improving and the corpus growing looked identical in every stored series. - lane:benchmark-capture runs one lane as a dry run over a fixed scope and stores every page it fetched in lane_benchmark_pages, which has no TTL, along with the refusal labels that applied at capture time - a replay mode in snapshotCache serves only benchmark pages, and an axios interceptor blocks the network, so an unseen page is a counted miss rather than a fetch - lane:scorecard replays every benchmark and stores one lane_scorecard_snapshots row per benchmark: planned values per field, the population a refusal could judge, the known-wrong count, replay coverage, and an order-independent fingerprint of the output - the lane-scorecard Development sweep stage runs it every sweep - the new collections are environment-local and never copied LLM lanes are excluded, because one run of an LLM lane is not repeatable. Refs #3526
quntao-z
force-pushed
the
feat/lane-benchmark-scorecard
branch
from
September 26, 2026 07:06
f74ec28 to
2dbb9a2
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Give the engine a thermometer: score each deterministic scraper lane against a frozen benchmark (pinned pages plus pinned refusal labels, replayed with the network blocked) and store one scorecard row per benchmark per sweep, so a lane's known-wrong count trends only when lane code changes. It must never be called precision, because a refusal is a negative label only. PR body: 'Refs #3526', no closing keyword, and state that benchmarks must be captured on Development after merge.
What Changed
lane:benchmark-capture, which runs a deterministic lane (BENCHMARKABLE_LANES) as a dry run and freezes both the pages it fetched (lane_benchmark_pages, no TTL) and the live refusals on rows it planned values for (lane_benchmarks). It also addslane:scorecard, which replays each benchmark through a new snapshot-cache benchmark mode. During replay the cache serves only pinned pages, the default axios instance is blocked, and the SSRF guard skips DNS lookups. The replay stores onelane_scorecard_snapshotsrow per benchmark withemitted,byField,labeledEntityEmitted,knownWrong,pagesMissed, and an order-independentoutputFingerprint. Because a refusal is a negative label only, the score isknownWrong / labeledEntityEmittedand is never reported as precision.lane-scorecardDevelopment sweep stage that runs the apply form every sweep. Benchmark runs set a newbenchmarkRunoption, so the orchestrator creates theirscrape_runsrows asinvalidated. Source health, freshness, and the barren-streak guard therefore never treat a replay as a live run. The three new collections are added toNEVER_COPY_COLLECTIONSand to the beta-to-development exclusions, so they stay environment-local.observationAssertsRefusedValuehelper so capture and lane attribution share it. Addsdocs/lane-scorecard.mdand links it fromAGENTS.mdanddocs/research-data-pipeline.md. Adds tests for benchmark cache mode, scorecard core, the orchestrator'sinvalidatedflag, and the sweep stage.This change only adds tooling. No benchmark exists until one is captured, so benchmarks must be captured on Development after merge with
yarn --cwd server lane:benchmark-capture ... --apply --confirm-lane-benchmark-capture. Until then, the sweep stage has nothing to score.Refs #3526
Risk Assessment
Testing
I ran the capture and scorecard CLIs as an operator would, against a local Mongo with live captures from medicine.yale.edu. The runs covered: - two identical apply sweeps, one of them with the network dead; - the missing-page case that failed in round 1, which now recordspagesMissed: 1andnetworkBlocks: 1, exits 0, and still stores the other benchmark's unchanged row; - frozen labels, whereknownWrongstayed at 2 after the live refusals were cleared; - a temporary lane-code edit, which changed the fingerprint and restored it when reverted, while dry runs stored nothing; - source health, which ignored all 15 invalidated benchmark runs. The focused unit tests also pass. There is no UI surface, so the evidence is CLI transcripts and persisted database state. The temporary Mongo, scripts and the scraper edit were removed, and the worktree is clean. The PR-body requirements ('Refs #3526', no closing keyword, capture benchmarks on Development after merge) belong to the PR phase, so I did not test them here.Evidence: Two replay sweeps, the second with network dead: identical stored rows
Evidence: Control: live fetch fails with the dead proxy
Evidence: Missing page: counted miss, exit 0, both benchmarks stored
Evidence: Frozen labels on the captured benchmark
Evidence: knownWrong unchanged after the live refusal was cleared
Evidence: Fingerprint moves with a lane code change and returns on revert
Evidence: Source health ignores benchmark scrape runs
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
server/src/scripts/laneBenchmarkRun.ts:59- Every replay (and capture) goes throughbuildOrchestrator().run, which creates a realScrapeRunrow for the lane (orchestrator.ts:68) with status success.buildSourceHealthRows(services/sourceHealthService.ts:273) takes the newest non-invalidated run per source aslatestRun, and nothing filters out dry runs. Thelane-scorecardstage runs after the live scrapes in every Development sweep, so each benchmarked lane's newest run is now a benchmark replay. Concrete sequence: the livedept-faculty-rosterrun fails or records materializationErrors, then the replay dry run succeeds a few minutes later. Source health readsok, andderivePromotionStatus(adminOperatorBoardService.ts:444) stops counting it as risk. That is exactly the fix(ops): Production's scheduled scrape crons show no evidence of ever having run #2513 failure mode, a green signal from a run that did not happen. The same rows also feedreadPriorRunYieldFacts/ the barren-streak guard (sourceYieldGuard.ts:104), where a--limit-scoped replay that emits nothing counts as a barren run of the lane, and they are mirrored to Productionscrape_runs. The stage comment at runScraperSweep.ts:1125-1127 ('writes only alane_scorecard_snapshotsrow') and docs/lane-scorecard.md are therefore wrong. Smallest fix: mark the benchmark run's ScrapeRuninvalidated: true(or give it a distinct triggeredBy and exclude that) right afterrunLaneDryreturns or throws. Health, freshness, barren-streak and retention readers already skip invalidated runs. The capture path in laneBenchmarkCapture.ts:113 needs the same fix.server/src/scrapers/snapshotBenchmarkMode.ts:79- Replay only blocks the default axios instance, but every benchmarkable lane callsassertPublicHttpUrlbeforegetCached(e.g. departmentRosterScraper.ts:3143, officialProfilePiBackfillScraper.ts:~2330, ysmFacultyDirectoryScraper.ts:~508, ysmAtoZScraper.ts:~130), and that runs a live DNS resolution (ssrfGuard.ts:282). So replay is not network-isolated. If a captured host stops resolving or starts resolving privately,SsrfBlockedErroris thrown before the cache read, the lane's per-target catch swallows it, and the output andknownWrongchange with no code change.pagesMisseddoes not count it, becausegetCachednever ran. This contradicts the intent that replay runs 'with the network blocked' and moves only when lane code changes. The remedy touches the SSRF guard or the lanes' fetch order, which is security-sensitive, so it needs a decision. Options: in replay mode, serve cached pages before the SSRF check, or record SSRF/DNS failures as counted misses.server/src/scripts/laneBenchmarkCapture.ts:145- Pages are inserted before theLaneBenchmarkrow is created. IfLaneBenchmark.createfails (e.g. a document too large because of labels, or a transient error), orphan pages remain. Retrying the same--idthen passes theLaneBenchmark.existscheck but fails on the unique{benchmarkId, sourceName, requestKey}index, so that id cannot be captured again without manual cleanup. Fix: create the benchmark row last but delete the pages for the id on failure, or checkLaneBenchmarkPage.existstoo.server/src/scrapers/__tests__/snapshotBenchmarkMode.test.ts:64- The test is titled 'blocks the network during replay and restores it afterwards', but its restore assertionexpect(axios.interceptors.request).toBeDefined()is always true, so it does not check the interceptor was ejected. Assert instead that a request afterfinishBenchmarkReplay()is not rejected withBenchmarkReplayNetworkError, for example with a stub adapter.🔧 Fix applied.
2 issues (1 warning, 1 info) still open:
server/src/scrapers/orchestrator.ts:75- Round 1 fixed review-1 by creating every benchmark ScrapeRun withinvalidated: true. That flag is also the operator quarantine fence.invalidatedScrapeRunIds(scrapers/invalidatedScrapeRuns.ts:28) loads every invalidated run id and refreshes every 30s. Its design note assumes the set is 'tiny and operator-driven (8 rows)'.isScrapeRunInvalidatedthen does anArray.includesover that list per entity key on the materializer write path (entityMaterializer.ts:5427, 6081). The lane-scorecard stage now adds one invalidated row per benchmark per Development sweep, and the capture CLI adds one per capture. These rows are never pruned, so the list grows without bound and every materialize lookup gets linearly slower. Benchmark runs are dry runs that emit no observations, so they never need quarantining. Smallest fix: keep the health/freshness/barren-streak exclusion, and exclude benchmark runs from the quarantine query (ScrapeRun.find({ invalidated: true, 'options.benchmarkRun': { $ne: true } })). The alternative is building a Set rather than callingincludes.server/src/services/adminOperatorBoardService.ts:438- This is a sibling reader the round 1 invalidation fix left behind.summarizeDryRunPosturedoes not skip invalidated runs.buildSourceFreshnesspasses it every run from the last 30 days. After each sweep, the operator board'slatestRunSummary.latestDryRuntherefore shows a lane-scorecard benchmark replay instead of the operator's latest real dry run. That contradicts the new docs claim that no freshness reader treats a benchmark run as a live run. Fix: filter outrun.invalidated(orrun.options?.benchmarkRun) insidesummarizeDryRunPosturebefore sorting.🔧 **Test** - 2 issues found → auto-fixed ✅
server/src/scripts/laneScorecard.ts:90- One page missing from a benchmark makeslane:scorecardexit 1 and store no row for any benchmark. Reproduced live: I removed one lab-homepage page from benchmarkatoz-test-1. Theysm-atoz-indexlane does not catch theBenchmarkReplayMissErrorthatgetCachednow throws, so it propagates throughrunLaneDryand the benchmark loop has no per-benchmark catch. The healthy benchmarkatoz-test-2then got 0 rows, but 1 row when run alone. This contradicts docs/lane-scorecard.md ('a page the capture never saw is a counted miss', 'pagesMissedis where that drift shows'). It also breaks the intent of one row per benchmark per sweep. A lane code change that fetches a page the capture never saw is exactly when the trend should record, yet it would fail the Development sweep'slane-scorecardstage on every run. Possible fixes: catch per benchmark inlaneScorecard.tsand record the failure plus the replay miss counts (stored row or report entry) while the other benchmarks continue, or have the miss return null so the lane's normal miss path runs andpagesMissedcounts it.Started an isolated MongoMemoryServer (mongod 8.2.6) on 127.0.0.1:27999, then rantsx src/scripts/buildMongoIndexes.ts --applyandtsx src/scrapers/seedSources.ts --apply --confirm-seed-applyagainst itSeeded 2 synthetic research_entities rows withfieldValueRefusals: one live websiteUrl refusal, one live name refusal, and one withdrawn refusaltsx src/scripts/laneBenchmarkCapture.ts --source=ysm-atoz-index --limit=2 --id=atoz-test-1(dry run, live network)tsx src/scripts/laneBenchmarkCapture.ts ... --applywithout the confirm flag, then with--confirm-lane-benchmark-captureRe-ran the capture with the same--idto check the frozen-benchmark refusaltsx src/scripts/laneScorecard.ts(dry run), then--applywithout and with--confirm-lane-scorecardRanlaneScorecard.ts --applyagain with a preloaded synthetic DNS outage (NODE_OPTIONS=--require /tmp/nm-break-dns.cjs), plus a control showing the outage makesassertPublicHttpUrlfail outside replayRemoved one benchmark page, then ranlaneScorecard.tsalone and alongside a second healthy benchmarkatoz-test-2Temporarily edited theysm-atoz-indexwebsiteUrl emission, replayed, reverted, and replayed againRanbuildSourceHealthRowsover the persistedscrape_runsrowsnpx vitest run src/scrapers/__tests__/snapshotBenchmarkMode.test.ts src/scripts/__tests__/laneScorecardCore.test.ts src/scrapers/__tests__/orchestrator.test.ts src/scripts/__tests__/runScraperSweep.test.ts🔧 Fix applied.
✅ Re-checked - no issues remain.
mongod(cached mongodb-memory-server binary) on 127.0.0.1:27017 dbylabs_lane_test, seeded withtsx src/scrapers/seedSources.ts --apply --confirm-seed-applytsx src/scripts/laneBenchmarkCapture.ts --source=ysm-atoz-index --limit=2 --id=atoz-r2-a --apply --confirm-lane-benchmark-capture(and --limit=3 --id=atoz-r2-b), live fetchtsx src/scripts/laneScorecard.ts --apply --confirm-lane-scorecardtwice, the second with HTTP(S)_PROXY pointed at a dead portControl: capture dry run with the dead proxy fails with ECONNREFUSED, proving the proxy blocks real networkAdversarial: deleted onelab-homepagepage fromatoz-r2-ainlane_benchmark_pages, then re-ranlane:scorecard --applySeeded syntheticfieldValueRefusals(one matching websiteUrl, one never-emitted shortDescription) viaplanFieldValueRefusal, capturedatoz-r2-labeled, scored it, then cleared the live refusals and re-scoredTemporarily editedysmAtoZScraper.ts:545output, ranlane:scorecard --benchmark=atoz-r2-bdry run, reverted the edit, and re-ranReadscrape_runsthroughbuildSourceHealthRowsto check benchmark runs are excludedvitest run src/scrapers/__tests__/snapshotBenchmarkMode.test.ts src/scripts/__tests__/laneScorecardCore.test.ts src/scrapers/__tests__/orchestrator.test.tsdocs/research-data-pipeline.md:94- The numbered Development post-run stage list in docs/research-data-pipeline.md is a hand-kept copy of DEVELOPMENT_POST_RUN_STAGE_DEFINITIONS in runScraperSweep.ts. I added the newlane-scorecardstage. The list was already missing profile-link-health, dead-research-website-clear, organization-identity-website-retire, and refusal-lane-attribution before this change. Follow-up: shorten the list to a pointer to the registry, or check it for drift against the registry, rather than keeping a prose copy in sync by hand.🔧 **Lint** - 1 issue found → no changes applied ✅
tsc --noEmitpasses. ESLint on the changed server files still needs a run afteryarn install:all.🔧 No changes applied.
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.