Skip to content

feat(extension): supplement Jev shortlist with BM25 candidate - #2960

Open
antonvishal wants to merge 2 commits into
browserbase:jev/2-decision-libraryfrom
antonvishal:bm25-candidate-eval
Open

antonvishal wants to merge 2 commits into
browserbase:jev/2-decision-libraryfrom
antonvishal:bm25-candidate-eval

Conversation

@antonvishal

@antonvishal antonvishal commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The existing implementation tries a lexical top-30 shortlist before asking Jev about the full candidate view. On lists with more than 40 candidates, repeated generic terms can fill those 30 slots and exclude a rare, relevant row.

This change keeps every candidate selected by the existing scorer. When all 30 slots are filled, it adds BM25's best candidate only if that candidate is outside the current shortlist. The list therefore contains at most 31 elements. If fewer than 30 candidates have lexical overlap, the current shortlist is returned without running BM25. Existing role views, tokenization, field weights, exact-name handling, and full-list retry remain in place. This may help when a rare, relevant term is missing from the existing shortlist while preserving its current candidates.

Measurements

Generated case Candidates Current top 30 BM25 top 30 Combined list
Invoice 1042 among 80 rows 80 rank 1 rank 1 included
Acme after 60 generic Account rows 81 absent rank 1 included, 31 options
Red toaster in a long product row 80 rank 1 rank 1 included
First record with longer context 80 rank 1 absent included, 30 options
"Erase the obsolete entry" against Archive/Delete labels 80 no lexical shortlist no lexical shortlist full-list path

Real websites

Stagehand captured accessibility snapshots in a local browser. Sites can change between runs, so counts below describe one capture.

Site AX nodes Pointer candidates Labeled target Current rank BM25 rank
Wikipedia population list 9,620 1,758 first India link 1 2
GitHub Stagehand issues 327 75 New issue button 1 1
Hacker News 756 231 first comments link 29 1
Books to Scrape 474 134 Sapiens Add to basket button 1 1

For 30 Hacker News "hide" links labeled by author and 20 Books to Scrape basket buttons labeled by title, the correct target was in the top 30 for 50/50 with the current scorer, 50/50 with BM25, and 50/50 with the combined list. BM25 reordered the first Hacker News comments link, but both scorers already included it. Jev selected the intended GitHub New issue and Sapiens basket controls from real snapshots; no controls were clicked.

The real website sample shows equal recall and provides no evidence of an accuracy gain. The Acme improvement and first-record regression come from deliberately constructed cases. One added option can still change Jev's answer on an untested page. A larger labeled set with complete act results is needed before claiming a general success-rate or latency improvement.

Verification

  • Extension unit suite: 49 files passed, 376 tests passed, 10 todo.
  • Extension TypeScript check, formatting, lint, and production Vite build passed.
  • Go embedded-extension archive regenerated; extensionpack --check passed.

Summary by cubic

Supplements Jev's lexical top-30 shortlist with BM25's best candidate so a rare, relevant row isn't excluded when repeated generic terms fill all 30 slots.

  • Adds scoreCandidatesBm25 and shortlistWithBm25 in bm25.ts; the scorer has no new dependency.
  • Keeps every existing shortlist candidate and appends BM25's top pick only when all 30 slots are filled and that pick is outside the shortlist, capping the list at 31.
  • Returns the unchanged shortlist without running BM25 when fewer than 30 candidates have lexical overlap.
  • Exports flattenText and tokens from tree.ts; pick.ts now uses the new helper, and existing role views, tokenization, field weights, exact-name handling, and full-list retry are untouched.
  • Adds unit tests for the rare-term-inclusion and first-record cases and regenerates the embedded extension archive.

Measurements

  • Real-site snapshots show equal recall among the current scorer, BM25, and the combined list, so no accuracy gain is demonstrated.
  • The Acme improvement and the first-record regression come from deliberately constructed cases; a larger labeled set with complete act results is needed before claiming a general gain.

Written for commit 6cadab8. Summary will update on new commits.

Review in cubic

@changeset-bot

changeset-bot Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 6cadab8

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

Copy link
Copy Markdown
Contributor

This PR is from an external contributor and must be approved by a stagehand team member with write access before CI can run.
Approving the latest commit mirrors it into an internal PR owned by the approver.
If new commits are pushed later, the internal PR stays open but is marked stale until someone approves the latest external commit and refreshes it.

@github-actions github-actions Bot added external-contributor Tracks PRs mirrored from external contributor forks. external-contributor:awaiting-approval Waiting for a stagehand team member to approve the latest external commit. labels Sep 17, 2026
@antonvishal
antonvishal marked this pull request as ready for review September 17, 2026 11:23
@antonvishal

Copy link
Copy Markdown
Contributor Author

@miguelg719 worth taking a look at BM25 when you pick those draft PRs for Jev. #2952

@miguelg719
miguelg719 force-pushed the jev/2-decision-library branch 2 times, most recently from 361fb16 to 547c6bc Compare September 21, 2026 21:35
@miguelg719
miguelg719 force-pushed the jev/2-decision-library branch from 547c6bc to 63a207a Compare September 21, 2026 21:45

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

external-contributor:awaiting-approval Waiting for a stagehand team member to approve the latest external commit. external-contributor Tracks PRs mirrored from external contributor forks.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants