Skip to content

feat(server): score scraper lanes on frozen benchmarks and store the trend - #3535

Merged
quntao-z merged 5 commits into
betafrom
feat/lane-benchmark-scorecard
Sep 26, 2026
Merged

quntao-z merged 5 commits into
betafrom
feat/lane-benchmark-scorecard

Conversation

@quntao-z

Copy link
Copy Markdown
Collaborator

Intent

Give the engine a thermometer: score each deterministic scraper lane against a frozen benchmark (pinned pages plus pinned refusal labels, replayed with the network blocked) and store one scorecard row per benchmark per sweep, so a lane's known-wrong count trends only when lane code changes. It must never be called precision, because a refusal is a negative label only. PR body: 'Refs #3526', no closing keyword, and state that benchmarks must be captured on Development after merge.

What Changed

  • Adds lane:benchmark-capture, which runs a deterministic lane (BENCHMARKABLE_LANES) as a dry run and freezes both the pages it fetched (lane_benchmark_pages, no TTL) and the live refusals on rows it planned values for (lane_benchmarks). It also adds lane:scorecard, which replays each benchmark through a new snapshot-cache benchmark mode. During replay the cache serves only pinned pages, the default axios instance is blocked, and the SSRF guard skips DNS lookups. The replay stores one lane_scorecard_snapshots row per benchmark with emitted, byField, labeledEntityEmitted, knownWrong, pagesMissed, and an order-independent outputFingerprint. Because a refusal is a negative label only, the score is knownWrong / labeledEntityEmitted and is never reported as precision.
  • Adds a lane-scorecard Development sweep stage that runs the apply form every sweep. Benchmark runs set a new benchmarkRun option, so the orchestrator creates their scrape_runs rows as invalidated. Source health, freshness, and the barren-streak guard therefore never treat a replay as a live run. The three new collections are added to NEVER_COPY_COLLECTIONS and to the beta-to-development exclusions, so they stay environment-local.
  • Moves the refusal-matching logic into an exported observationAssertsRefusedValue helper so capture and lane attribution share it. Adds docs/lane-scorecard.md and links it from AGENTS.md and docs/research-data-pipeline.md. Adds tests for benchmark cache mode, scorecard core, the orchestrator's invalidated flag, and the sweep stage.

This change only adds tooling. No benchmark exists until one is captured, so benchmarks must be captured on Development after merge with yarn --cwd server lane:benchmark-capture ... --apply --confirm-lane-benchmark-capture. Until then, the sweep stage has nothing to score.

Refs #3526

Risk Assessment

⚠️ Medium: The round 1 fixes work, but they reuse the operator quarantine flag, which grows the materializer's invalidated-run set without bound and still leaves one operator-board reader showing benchmark replays; both are follow-up fixes and neither corrupts data.

Testing

I ran the capture and scorecard CLIs as an operator would, against a local Mongo with live captures from medicine.yale.edu. The runs covered: - two identical apply sweeps, one of them with the network dead; - the missing-page case that failed in round 1, which now records pagesMissed: 1 and networkBlocks: 1, exits 0, and still stores the other benchmark's unchanged row; - frozen labels, where knownWrong stayed at 2 after the live refusals were cleared; - a temporary lane-code edit, which changed the fingerprint and restored it when reverted, while dry runs stored nothing; - source health, which ignored all 15 invalidated benchmark runs. The focused unit tests also pass. There is no UI surface, so the evidence is CLI transcripts and persisted database state. The temporary Mongo, scripts and the scraper edit were removed, and the worktree is clean. The PR-body requirements ('Refs #3526', no closing keyword, capture benchmarks on Development after merge) belong to the PR phase, so I did not test them here.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Operator captures a benchmark from a live lane run, freezing its pages ✅ pass live r2/01-capture-a.log, r2/02-capture-b.log (pageCount 3 and 4, apply)
Replay with the network blocked stores one row per benchmark per sweep, and two sweeps give identical fingerprints ✅ pass live r2/03, r2/04 (dead proxy), r2/05-stored-rows.txt, control r2/06
Adversarial: a page missing from one benchmark is a counted miss, the sweep succeeds, and the other benchmark still gets its row ✅ pass live r2/08-scorecard-apply-missing-page.log: rc=0, stored 2, atoz-r2-a pagesMissed 1 networkBlocks 1, atoz-r2-b fingerprint unchanged
Refusals are frozen as labels at capture and knownWrong does not move when live refusals change ✅ pass live r2/11-frozen-labels.txt, r2/14-labeled-rows.txt (knownWrong 2, labelsMatched 1 of 2, both before and after clearing)
knownWrong is scored over the labeled population and never labelled precision ✅ pass live r2/14 byField labeledEntityEmitted; no 'precision' in the diff's added lines or the CLI output
A lane code change moves the output fingerprint and reverting restores it, while dry runs store no row ✅ pass live r2/17-lane-change-comparison.txt; row count 8 matches the apply runs only
Benchmark scrape runs never read as live runs of the lane in source health ✅ pass live r2/18-source-health-ignores-benchmark-runs.txt (15 runs, all invalidated, recentRuns total 0, latestRun null)
The PR body says 'Refs #3526' with no closing keyword and notes that benchmarks are captured on Development after merge ⏸️ untested no The PR is created by the outer executor's PR phase, not this test phase, so there is no PR body to inspect yet.
Evidence: Two replay sweeps, the second with network dead: identical stored rows
{
  benchmarkId: 'atoz-r2-a',
  codeSha: '54f0d063c4d86decdecdc697116f91023ebc32d1',
  pagesServed: 3,
  pagesMissed: 0,
  emitted: 41,
  knownWrong: 0,
  outputFingerprint: '9ab548edc1ee6209d271191a69b0d756df175ce41cb8b855b5dd01c751c9444a'
}
{
  benchmarkId: 'atoz-r2-b',
  codeSha: '54f0d063c4d86decdecdc697116f91023ebc32d1',
  pagesServed: 4,
  pagesMissed: 0,
  emitted: 48,
  knownWrong: 0,
  outputFingerprint: 'cb72aaf895744814ab4b6171d8866948b9af0e08649f2e2771aa063bbd628aad'
}
{
  benchmarkId: 'atoz-r2-a',
  codeSha: '54f0d063c4d86decdecdc697116f91023ebc32d1',
  pagesServed: 3,
  pagesMissed: 0,
  emitted: 41,
  knownWrong: 0,
  outputFingerprint: '9ab548edc1ee6209d271191a69b0d756df175ce41cb8b855b5dd01c751c9444a'
}
{
  benchmarkId: 'atoz-r2-b',
  codeSha: '54f0d063c4d86decdecdc697116f91023ebc32d1',
  pagesServed: 4,
  pagesMissed: 0,
  emitted: 48,
  knownWrong: 0,
  outputFingerprint: 'cb72aaf895744814ab4b6171d8866948b9af0e08649f2e2771aa063bbd628aad'
}
Evidence: Control: live fetch fails with the dead proxy
Environment: development; Mongo target: 127.0.0.1/ylabs_lane_test; mode: dry-run
(node:60785) [DEP0205] DeprecationWarning: `module.register()` is deprecated. Use `module.registerHooks()` instead.
(Use `node --trace-deprecation ...` to show where the warning was created)
Connected to database 🚀
MongoDB: 5 declared index(es) are missing across 3 collection(s). Indexes are no longer built on connect; run `yarn --cwd server db:build-indexes --apply`.
MongoDB: lane_benchmarks missing benchmarkId_1
MongoDB: lane_benchmark_pages missing benchmarkId_1_sourceName_1_requestKey_1
MongoDB: scrape_runs missing sourceId_1_startedAt_-1, status_1_startedAt_-1, invalidated_1
[ysm-atoz-index] Fetching https://medicine.yale.edu/about/a-to-z-index/lab-websites/
Error: connect ECONNREFUSED 127.0.0.1:9
Error: connect ECONNREFUSED 127.0.0.1:9
    at AxiosError.from (file://~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/node_modules/axios/lib/core/AxiosError.js:77:24)
    at RedirectableRequest.handleRequestError (file://~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/node_modules/axios/lib/adapters/http.js:1259:27)
    at RedirectableRequest.emit (node:events:526:24)
    at eventHandlers.<computed> (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/node_modules/follow-redirects/index.js:56:24)
    at ClientRequest.emit (node:events:514:20)
    at onerror (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/node_modules/axios/node_modules/agent-base/src/index.ts:236:9)
    at callbackError (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/node_modules/axios/node_modules/agent-base/src/index.ts:258:5)
    at process.processTicksAndRejections (node:internal/process/task_queues:104:5)
    at Axios.request (file://~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/node_modules/axios/lib/core/Axios.js:46:41)
    at process.processTicksAndRejections (node:internal/process/task_queues:104:5)
    at async fetchPage (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/src/scrapers/sources/ysmAtoZScraper.ts:133:15)
    at async YsmAtoZScraper.run (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/src/scrapers/sources/ysmAtoZScraper.ts:684:18)
    at async ScraperOrchestrator.run (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/src/scrapers/orchestrator.ts:133:23)
    at async runLaneDry (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/src/scripts/laneBenchmarkRun.ts:60:55)
    at async main (~/.no-mistakes/worktrees/07c70b5bbe11/01M3E4WB19R2PHA35NV93MGR2M/server/src/scripts/laneBenchmarkCapture.ts:114:11)
Evidence: Missing page: counted miss, exit 0, both benchmarks stored
Environment: development; Mongo target: 127.0.0.1/ylabs_lane_test; mode: apply
(node:60858) [DEP0205] DeprecationWarning: `module.register()` is deprecated. Use `module.registerHooks()` instead.
(Use `node --trace-deprecation ...` to show where the warning was created)
Connected to database 🚀
MongoDB: 6 declared index(es) are missing across 4 collection(s). Indexes are no longer built on connect; run `yarn --cwd server db:build-indexes --apply`.
MongoDB: lane_benchmarks missing benchmarkId_1
MongoDB: lane_benchmark_pages missing benchmarkId_1_sourceName_1_requestKey_1
MongoDB: lane_scorecard_snapshots missing benchmarkId_1_measuredAt_-1
MongoDB: scrape_runs missing sourceId_1_startedAt_-1, status_1_startedAt_-1, invalidated_1
[ysm-atoz-index] Fetching https://medicine.yale.edu/about/a-to-z-index/lab-websites/
[ysm-atoz-index] Parsed 262 labs from index
[observation-store] ysm-atoz-index asserted 1 observation(s) on a retired field; nothing reads them, so they are refused at ingest (#3362).
[observation-store] ysm-atoz-index asserted 1 observation(s) on a retired field; nothing reads them, so they are refused at ingest (#3362).
[ysm-atoz-index] A-Z index snapshot marked incomplete (parsed=262, narrowed=true); delisting detection will not act on this run
[ysm-atoz-index] Emitted 37 observations across 2 labs
[ysm-atoz-index] Inferred PI for 0/2 labs
[ysm-atoz-index] Found official homepage descriptions for 1/2 labs
[ysm-atoz-index] Fetching https://medicine.yale.edu/about/a-to-z-index/lab-websites/
[ysm-atoz-index] Parsed 262 labs from index
[observation-store] ysm-atoz-index asserted 1 observation(s) on a retired field; nothing reads them, so they are refused at ingest (#3362).
[observation-store] ysm-atoz-index asserted 1 observation(s) on a retired field; nothing reads them, so they are refused at ingest (#3362).
[observation-store] ysm-atoz-index asserted 1 observation(s) on a retired field; nothing reads them, so they are refused at ingest (#3362).
[ysm-atoz-index] A-Z index snapshot marked incomplete (parsed=262, narrowed=true); delisting detection will not act on this run
[ysm-atoz-index] Emitted 48 observations across 3 labs
[ysm-atoz-index] Inferred PI for 0/3 labs
[ysm-atoz-index] Found official homepage descriptions for 2/3 labs
{
  "script": "lane:scorecard",
  "mode": "apply",
  "benchmarks": 2,
  "stored": 2,
  "results": [
    {
      "measuredAt": "2026-09-26T06:51:50.891Z",
      "environment": "development",
      "databaseName": "ylabs_lane_test",
      "benchmarkId": "atoz-r2-a",
      "sourceName": "ysm-atoz-index",
      "codeSha": "54f0d063c4d86decdecdc697116f91023ebc32d1",
      "pagesServed": 2,
      "pagesMissed": 1,
      "emitted": 37,
      "knownWrong": 0,
      "labelsMatched": 0,
      "labelCount": 0,
      "outputFingerprint": "9c0a274dec5137d74c9f1050970a2a903c6ef6b8fca101a72fae722e8c3b7ecd",
      "byField": [
        {
          "field": "contactEmail",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "contactName",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "contactRole",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "dataSources",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "displayName",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "email",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "entityType",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "fname",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "fullDescription",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "inferredPiUserKey",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "inferredUserName",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "kind",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "lname",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "name",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "netid",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "profileUrl",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "profileUrls",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "profileVerified",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "researchGroupKey",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "role",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "school",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "shortDescription",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "slug",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "sourceUrls",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "title",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "websiteUrl",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "ysmLabIndexHealth",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        }
      ],
      "networkBlocks": 1,
      "truncated": false
    },
    {
      "measuredAt": "2026-09-26T06:51:51.873Z",
      "environment": "development",
      "databaseName": "ylabs_lane_test",
      "benchmarkId": "atoz-r2-b",
      "sourceName": "ysm-atoz-index",
      "codeSha": "54f0d063c4d86decdecdc697116f91023ebc32d1",
      "pagesServed": 4,
      "pagesMissed": 0,
      "emitted": 48,
      "knownWrong": 0,
      "labelsMatched": 0,
      "labelCount": 0,
      "outputFingerprint": "cb72aaf895744814ab4b6171d8866948b9af0e08649f2e2771aa063bbd628aad",
      "byField": [
        {
          "field": "contactEmail",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "contactName",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "contactRole",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "dataSources",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "displayName",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "email",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "entityType",
          "emitted": 5,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "fname",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "fullDescription",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "inferredPiUserKey",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "inferredUserName",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "kind",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "lname",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "name",
          "emitted": 4,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "netid",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "profileUrl",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "profileUrls",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "profileVerified",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "researchGroupKey",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "role",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "school",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "shortDescription",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "slug",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "sourceUrls",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "title",
          "emitted": 2,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "websiteUrl",
          "emitted": 3,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        },
        {
          "field": "ysmLabIndexHealth",
          "emitted": 1,
          "labeledEntityEmitted": 0,
          "knownWrong": 0
        }
      ],
      "networkBlocks": 0,
      "truncated": false
    }
  ]
}
MongoDB: disconnected
rc=0
Evidence: Frozen labels on the captured benchmark
{
  pageCount: 4,
  labels: [
    {
      entityKey: 'ysm-3d-tumor-lab',
      field: 'websiteUrl',
      valueKey: 'medicine.yale.edu/lab/3d-tumor-lab',
      rule: 'wrong_owner'
    },
    {
      entityKey: 'ysm-3d-tumor-lab',
      field: 'shortDescription',
      valueKey: 'a synthetic description this lane never emits',
      rule: 'operator_judgement'
    }
  ]
}
Evidence: knownWrong unchanged after the live refusal was cleared
{
  measuredAt: ISODate('2026-09-26T06:53:59.021Z'),
  emitted: 48,
  knownWrong: 2,
  labelsMatched: 1,
  labelCount: 2,
  outputFingerprint: 'cb72aaf895744814ab4b6171d8866948b9af0e08649f2e2771aa063bbd628aad',
  byFieldLabeled: [
    {
      field: 'shortDescription',
      emitted: 2,
      labeledEntityEmitted: 1,
      knownWrong: 0
    },
    {
      field: 'sourceUrls',
      emitted: 3,
      labeledEntityEmitted: 1,
      knownWrong: 1
    },
    {
      field: 'websiteUrl',
      emitted: 3,
      labeledEntityEmitted: 1,
      knownWrong: 1
    }
  ]
}
{
  measuredAt: ISODate('2026-09-26T06:54:09.036Z'),
  emitted: 48,
  knownWrong: 2,
  labelsMatched: 1,
  labelCount: 2,
  outputFingerprint: 'cb72aaf895744814ab4b6171d8866948b9af0e08649f2e2771aa063bbd628aad',
  byFieldLabeled: [
    {
      field: 'shortDescription',
      emitted: 2,
      labeledEntityEmitted: 1,
      knownWrong: 0
    },
    {
      field: 'sourceUrls',
      emitted: 3,
      labeledEntityEmitted: 1,
      knownWrong: 1
    },
    {
      field: 'websiteUrl',
      emitted: 3,
      labeledEntityEmitted: 1,
      knownWrong: 1
    }
  ]
}
Evidence: Fingerprint moves with a lane code change and returns on revert
15-dryrun-after-lane-change:       "outputFingerprint": "ae9577f1069a2031eacabba0438e09097f15b9a4da8d6cdbc65fdbe2596f8fef",   "stored": 0,
16-dryrun-after-revert:       "outputFingerprint": "cb72aaf895744814ab4b6171d8866948b9af0e08649f2e2771aa063bbd628aad",   "stored": 0,
baseline apply rows atoz-r2-b fingerprint: cb72aaf8...aad
Evidence: Source health ignores benchmark scrape runs
{"atozRunsInDb":15,"allInvalidated":true,"allFlaggedBenchmark":true,"healthRecentRuns":{"total":0,"success":0,"partial":0,"failure":0,"running":0},"healthLatestRun":null,"risk":"warn"}
- Outcome: 🔧 2 issues found → auto-fixed ✅ across 2 runs (35m5s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 issues (1 warning, 1 info)
  • 🚨 server/src/scripts/laneBenchmarkRun.ts:59 - Every replay (and capture) goes through buildOrchestrator().run, which creates a real ScrapeRun row for the lane (orchestrator.ts:68) with status success. buildSourceHealthRows (services/sourceHealthService.ts:273) takes the newest non-invalidated run per source as latestRun, and nothing filters out dry runs. The lane-scorecard stage runs after the live scrapes in every Development sweep, so each benchmarked lane's newest run is now a benchmark replay. Concrete sequence: the live dept-faculty-roster run fails or records materializationErrors, then the replay dry run succeeds a few minutes later. Source health reads ok, and derivePromotionStatus (adminOperatorBoardService.ts:444) stops counting it as risk. That is exactly the fix(ops): Production's scheduled scrape crons show no evidence of ever having run #2513 failure mode, a green signal from a run that did not happen. The same rows also feed readPriorRunYieldFacts / the barren-streak guard (sourceYieldGuard.ts:104), where a --limit-scoped replay that emits nothing counts as a barren run of the lane, and they are mirrored to Production scrape_runs. The stage comment at runScraperSweep.ts:1125-1127 ('writes only a lane_scorecard_snapshots row') and docs/lane-scorecard.md are therefore wrong. Smallest fix: mark the benchmark run's ScrapeRun invalidated: true (or give it a distinct triggeredBy and exclude that) right after runLaneDry returns or throws. Health, freshness, barren-streak and retention readers already skip invalidated runs. The capture path in laneBenchmarkCapture.ts:113 needs the same fix.
  • ⚠️ server/src/scrapers/snapshotBenchmarkMode.ts:79 - Replay only blocks the default axios instance, but every benchmarkable lane calls assertPublicHttpUrl before getCached (e.g. departmentRosterScraper.ts:3143, officialProfilePiBackfillScraper.ts:~2330, ysmFacultyDirectoryScraper.ts:~508, ysmAtoZScraper.ts:~130), and that runs a live DNS resolution (ssrfGuard.ts:282). So replay is not network-isolated. If a captured host stops resolving or starts resolving privately, SsrfBlockedError is thrown before the cache read, the lane's per-target catch swallows it, and the output and knownWrong change with no code change. pagesMissed does not count it, because getCached never ran. This contradicts the intent that replay runs 'with the network blocked' and moves only when lane code changes. The remedy touches the SSRF guard or the lanes' fetch order, which is security-sensitive, so it needs a decision. Options: in replay mode, serve cached pages before the SSRF check, or record SSRF/DNS failures as counted misses.
  • ℹ️ server/src/scripts/laneBenchmarkCapture.ts:145 - Pages are inserted before the LaneBenchmark row is created. If LaneBenchmark.create fails (e.g. a document too large because of labels, or a transient error), orphan pages remain. Retrying the same --id then passes the LaneBenchmark.exists check but fails on the unique {benchmarkId, sourceName, requestKey} index, so that id cannot be captured again without manual cleanup. Fix: create the benchmark row last but delete the pages for the id on failure, or check LaneBenchmarkPage.exists too.
  • ℹ️ server/src/scrapers/__tests__/snapshotBenchmarkMode.test.ts:64 - The test is titled 'blocks the network during replay and restores it afterwards', but its restore assertion expect(axios.interceptors.request).toBeDefined() is always true, so it does not check the interceptor was ejected. Assert instead that a request after finishBenchmarkReplay() is not rejected with BenchmarkReplayNetworkError, for example with a stub adapter.

🔧 Fix applied.
2 issues (1 warning, 1 info) still open:

  • ⚠️ server/src/scrapers/orchestrator.ts:75 - Round 1 fixed review-1 by creating every benchmark ScrapeRun with invalidated: true. That flag is also the operator quarantine fence. invalidatedScrapeRunIds (scrapers/invalidatedScrapeRuns.ts:28) loads every invalidated run id and refreshes every 30s. Its design note assumes the set is 'tiny and operator-driven (8 rows)'. isScrapeRunInvalidated then does an Array.includes over that list per entity key on the materializer write path (entityMaterializer.ts:5427, 6081). The lane-scorecard stage now adds one invalidated row per benchmark per Development sweep, and the capture CLI adds one per capture. These rows are never pruned, so the list grows without bound and every materialize lookup gets linearly slower. Benchmark runs are dry runs that emit no observations, so they never need quarantining. Smallest fix: keep the health/freshness/barren-streak exclusion, and exclude benchmark runs from the quarantine query (ScrapeRun.find({ invalidated: true, &#39;options.benchmarkRun&#39;: { $ne: true } })). The alternative is building a Set rather than calling includes.
  • ℹ️ server/src/services/adminOperatorBoardService.ts:438 - This is a sibling reader the round 1 invalidation fix left behind. summarizeDryRunPosture does not skip invalidated runs. buildSourceFreshness passes it every run from the last 30 days. After each sweep, the operator board's latestRunSummary.latestDryRun therefore shows a lane-scorecard benchmark replay instead of the operator's latest real dry run. That contradicts the new docs claim that no freshness reader treats a benchmark run as a live run. Fix: filter out run.invalidated (or run.options?.benchmarkRun) inside summarizeDryRunPosture before sorting.
🔧 **Test** - 2 issues found → auto-fixed ✅
  • ⚠️ server/src/scripts/laneScorecard.ts:90 - One page missing from a benchmark makes lane:scorecard exit 1 and store no row for any benchmark. Reproduced live: I removed one lab-homepage page from benchmark atoz-test-1. The ysm-atoz-index lane does not catch the BenchmarkReplayMissError that getCached now throws, so it propagates through runLaneDry and the benchmark loop has no per-benchmark catch. The healthy benchmark atoz-test-2 then got 0 rows, but 1 row when run alone. This contradicts docs/lane-scorecard.md ('a page the capture never saw is a counted miss', 'pagesMissed is where that drift shows'). It also breaks the intent of one row per benchmark per sweep. A lane code change that fetches a page the capture never saw is exactly when the trend should record, yet it would fail the Development sweep's lane-scorecard stage on every run. Possible fixes: catch per benchmark in laneScorecard.ts and record the failure plus the replay miss counts (stored row or report entry) while the other benchmarks continue, or have the miss return null so the lane's normal miss path runs and pagesMissed counts it.
  • 🚨 live validation verdict: no-go (9 of 10 scenarios were driven live against the product); failed: Adversarial: a page missing from a benchmark is a counted miss, and every benchmark still gets its row
  • Live validation: ❌ no-go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Operator captures a benchmark in dry run: report shows pages, planned values and labels; nothing is persisted ✅ pass live 01-capture-dry-run.log (pageCount 3, plannedObservationCount 41; lane_benchmarks and lane_benchmark_pages both 0)
Operator applies a capture: pages and live refusal labels are frozen, withdrawn refusals are excluded, and --apply without the confirm flag is refused ✅ pass live 02-capture-apply-no-confirm.log, 03-capture-apply.log, 04-capture-persisted-state.txt (3 pages, 2 labels, withdrawn refusal excluded)
Adversarial: re-capturing an existing benchmark id is refused because a benchmark is frozen ✅ pass live 05-capture-refuses-existing-id.log
Operator runs the scorecard: a replay with the network blocked serves only captured pages and reports emitted, knownWrong, per-field labeledEntityEmitted and a fingerprint, never 'precision' ✅ pass live 06-scorecard-dry-run.log (pagesServed 3, pagesMissed 0, networkBlocks 0, knownWrong 2; grep finds no 'precision')
Scorecard --apply stores one row per benchmark per run and requires --confirm-lane-scorecard ✅ pass live 07-scorecard-apply-no-confirm.log, 10-stored-scorecard-rows.txt
Adversarial: replaying unchanged code twice, the second time under a DNS outage, gives identical stored numbers and fingerprint ✅ pass live 10-stored-scorecard-rows.txt (both rows fingerprint 9ab548ed...), 11-dns-break-control.txt (the outage breaks the SSRF DNS lookup outside replay)
Changing lane code moves knownWrong and the fingerprint; reverting restores them ✅ pass live 18-lane-change-comparison.txt
Benchmark capture and replay scrape_runs rows are invalidated, so source health never reads them as live runs; no scrape_snapshots or observations are written ✅ pass live 19-source-health-ignores-benchmark-runs.txt (11 runs, all invalidated, health latestRun null), 10-stored-scorecard-rows.txt (liveSnapshots 0, observations 0)
Adversarial: a page missing from a benchmark is a counted miss, and every benchmark still gets its row ❌ fail live 12-scorecard-missing-page.log (BenchmarkReplayMissError, exit 1), 14-summary.txt (healthy atoz-test-2 got 0 rows)
The Development sweep runs the lane-scorecard stage after the live scrapes ⏸️ untested no A full Development sweep needs Development MongoDB and Meilisearch credentials and runs every live scraper, so I could not stand it up in isolation. I drove the stage's exact command (`lane:scorecard…
  • Started an isolated MongoMemoryServer (mongod 8.2.6) on 127.0.0.1:27999, then ran tsx src/scripts/buildMongoIndexes.ts --apply and tsx src/scrapers/seedSources.ts --apply --confirm-seed-apply against it
  • Seeded 2 synthetic research_entities rows with fieldValueRefusals: one live websiteUrl refusal, one live name refusal, and one withdrawn refusal
  • tsx src/scripts/laneBenchmarkCapture.ts --source=ysm-atoz-index --limit=2 --id=atoz-test-1 (dry run, live network)
  • tsx src/scripts/laneBenchmarkCapture.ts ... --apply without the confirm flag, then with --confirm-lane-benchmark-capture
  • Re-ran the capture with the same --id to check the frozen-benchmark refusal
  • tsx src/scripts/laneScorecard.ts (dry run), then --apply without and with --confirm-lane-scorecard
  • Ran laneScorecard.ts --apply again with a preloaded synthetic DNS outage (NODE_OPTIONS=--require /tmp/nm-break-dns.cjs), plus a control showing the outage makes assertPublicHttpUrl fail outside replay
  • Removed one benchmark page, then ran laneScorecard.ts alone and alongside a second healthy benchmark atoz-test-2
  • Temporarily edited the ysm-atoz-index websiteUrl emission, replayed, reverted, and replayed again
  • Ran buildSourceHealthRows over the persisted scrape_runs rows
  • npx vitest run src/scrapers/__tests__/snapshotBenchmarkMode.test.ts src/scripts/__tests__/laneScorecardCore.test.ts src/scrapers/__tests__/orchestrator.test.ts src/scripts/__tests__/runScraperSweep.test.ts

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Operator captures a benchmark from a live lane run, freezing its pages ✅ pass live r2/01-capture-a.log, r2/02-capture-b.log (pageCount 3 and 4, apply)
Replay with the network blocked stores one row per benchmark per sweep, and two sweeps give identical fingerprints ✅ pass live r2/03, r2/04 (dead proxy), r2/05-stored-rows.txt, control r2/06
Adversarial: a page missing from one benchmark is a counted miss, the sweep succeeds, and the other benchmark still gets its row ✅ pass live r2/08-scorecard-apply-missing-page.log: rc=0, stored 2, atoz-r2-a pagesMissed 1 networkBlocks 1, atoz-r2-b fingerprint unchanged
Refusals are frozen as labels at capture and knownWrong does not move when live refusals change ✅ pass live r2/11-frozen-labels.txt, r2/14-labeled-rows.txt (knownWrong 2, labelsMatched 1 of 2, both before and after clearing)
knownWrong is scored over the labeled population and never labelled precision ✅ pass live r2/14 byField labeledEntityEmitted; no 'precision' in the diff's added lines or the CLI output
A lane code change moves the output fingerprint and reverting restores it, while dry runs store no row ✅ pass live r2/17-lane-change-comparison.txt; row count 8 matches the apply runs only
Benchmark scrape runs never read as live runs of the lane in source health ✅ pass live r2/18-source-health-ignores-benchmark-runs.txt (15 runs, all invalidated, recentRuns total 0, latestRun null)
The PR body says 'Refs #3526' with no closing keyword and notes that benchmarks are captured on Development after merge ⏸️ untested no The PR is created by the outer executor's PR phase, not this test phase, so there is no PR body to inspect yet.
  • mongod (cached mongodb-memory-server binary) on 127.0.0.1:27017 db ylabs_lane_test, seeded with tsx src/scrapers/seedSources.ts --apply --confirm-seed-apply
  • tsx src/scripts/laneBenchmarkCapture.ts --source=ysm-atoz-index --limit=2 --id=atoz-r2-a --apply --confirm-lane-benchmark-capture (and --limit=3 --id=atoz-r2-b), live fetch
  • tsx src/scripts/laneScorecard.ts --apply --confirm-lane-scorecard twice, the second with HTTP(S)_PROXY pointed at a dead port
  • Control: capture dry run with the dead proxy fails with ECONNREFUSED, proving the proxy blocks real network
  • Adversarial: deleted one lab-homepage page from atoz-r2-a in lane_benchmark_pages, then re-ran lane:scorecard --apply
  • Seeded synthetic fieldValueRefusals (one matching websiteUrl, one never-emitted shortDescription) via planFieldValueRefusal, captured atoz-r2-labeled, scored it, then cleared the live refusals and re-scored
  • Temporarily edited ysmAtoZScraper.ts:545 output, ran lane:scorecard --benchmark=atoz-r2-b dry run, reverted the edit, and re-ran
  • Read scrape_runs through buildSourceHealthRows to check benchmark runs are excluded
  • vitest run src/scrapers/__tests__/snapshotBenchmarkMode.test.ts src/scripts/__tests__/laneScorecardCore.test.ts src/scrapers/__tests__/orchestrator.test.ts
⚠️ **Document** - 1 info
  • ℹ️ docs/research-data-pipeline.md:94 - The numbered Development post-run stage list in docs/research-data-pipeline.md is a hand-kept copy of DEVELOPMENT_POST_RUN_STAGE_DEFINITIONS in runScraperSweep.ts. I added the new lane-scorecard stage. The list was already missing profile-link-health, dead-research-website-clear, organization-identity-website-retire, and refusal-lane-attribution before this change. Follow-up: shorten the list to a pointer to the registry, or check it for drift against the registry, rather than keeping a prose copy in sync by hand.
🔧 **Lint** - 1 issue found → no changes applied ✅
  • ⚠️ ESLint could not run: the worktree has no root node_modules, so eslint.config.js fails to import '@eslint/js'. Prettier (pinned 3.8.3) passes on every changed file, and server tsc --noEmit passes. ESLint on the changed server files still needs a run after yarn install:all.

🔧 No changes applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

…he trend

Nothing measured a lane against a fixed input, so a lane improving and
the corpus growing looked identical in every stored series.

- lane:benchmark-capture runs one lane as a dry run over a fixed scope
  and stores every page it fetched in lane_benchmark_pages, which has no
  TTL, along with the refusal labels that applied at capture time
- a replay mode in snapshotCache serves only benchmark pages, and an
  axios interceptor blocks the network, so an unseen page is a counted
  miss rather than a fetch
- lane:scorecard replays every benchmark and stores one
  lane_scorecard_snapshots row per benchmark: planned values per field,
  the population a refusal could judge, the known-wrong count, replay
  coverage, and an order-independent fingerprint of the output
- the lane-scorecard Development sweep stage runs it every sweep
- the new collections are environment-local and never copied

LLM lanes are excluded, because one run of an LLM lane is not repeatable.

Refs #3526
@quntao-z
quntao-z force-pushed the feat/lane-benchmark-scorecard branch from f74ec28 to 2dbb9a2 Compare September 26, 2026 07:06
@quntao-z
quntao-z merged commit dce1673 into beta Sep 26, 2026
2 checks passed
@quntao-z
quntao-z deleted the feat/lane-benchmark-scorecard branch September 26, 2026 07:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant