You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(extension): supplement Jev shortlist with BM25 candidate - #2960
The existing implementation tries a lexical top-30 shortlist before asking Jev about the full candidate view. On lists with more than 40 candidates, repeated generic terms can fill those 30 slots and exclude a rare, relevant row.
This change keeps every candidate selected by the existing scorer. When all 30 slots are filled, it adds BM25's best candidate only if that candidate is outside the current shortlist. The list therefore contains at most 31 elements. If fewer than 30 candidates have lexical overlap, the current shortlist is returned without running BM25. Existing role views, tokenization, field weights, exact-name handling, and full-list retry remain in place. This may help when a rare, relevant term is missing from the existing shortlist while preserving its current candidates.
Measurements
Generated case
Candidates
Current top 30
BM25 top 30
Combined list
Invoice 1042 among 80 rows
80
rank 1
rank 1
included
Acme after 60 generic Account rows
81
absent
rank 1
included, 31 options
Red toaster in a long product row
80
rank 1
rank 1
included
First record with longer context
80
rank 1
absent
included, 30 options
"Erase the obsolete entry" against Archive/Delete labels
80
no lexical shortlist
no lexical shortlist
full-list path
Real websites
Stagehand captured accessibility snapshots in a local browser. Sites can change between runs, so counts below describe one capture.
Site
AX nodes
Pointer candidates
Labeled target
Current rank
BM25 rank
Wikipedia population list
9,620
1,758
first India link
1
2
GitHub Stagehand issues
327
75
New issue button
1
1
Hacker News
756
231
first comments link
29
1
Books to Scrape
474
134
Sapiens Add to basket button
1
1
For 30 Hacker News "hide" links labeled by author and 20 Books to Scrape basket buttons labeled by title, the correct target was in the top 30 for 50/50 with the current scorer, 50/50 with BM25, and 50/50 with the combined list. BM25 reordered the first Hacker News comments link, but both scorers already included it. Jev selected the intended GitHub New issue and Sapiens basket controls from real snapshots; no controls were clicked.
The real website sample shows equal recall and provides no evidence of an accuracy gain. The Acme improvement and first-record regression come from deliberately constructed cases. One added option can still change Jev's answer on an untested page. A larger labeled set with complete act results is needed before claiming a general success-rate or latency improvement.
Extension TypeScript check, formatting, lint, and production Vite build passed.
Go embedded-extension archive regenerated; extensionpack --check passed.
Summary by cubic
Supplements Jev's lexical top-30 shortlist with BM25's best candidate so a rare, relevant row isn't excluded when repeated generic terms fill all 30 slots.
Adds scoreCandidatesBm25 and shortlistWithBm25 in bm25.ts; the scorer has no new dependency.
Keeps every existing shortlist candidate and appends BM25's top pick only when all 30 slots are filled and that pick is outside the shortlist, capping the list at 31.
Returns the unchanged shortlist without running BM25 when fewer than 30 candidates have lexical overlap.
Exports flattenText and tokens from tree.ts; pick.ts now uses the new helper, and existing role views, tokenization, field weights, exact-name handling, and full-list retry are untouched.
Adds unit tests for the rare-term-inclusion and first-record cases and regenerates the embedded extension archive.
Measurements
Real-site snapshots show equal recall among the current scorer, BM25, and the combined list, so no accuracy gain is demonstrated.
The Acme improvement and the first-record regression come from deliberately constructed cases; a larger labeled set with complete act results is needed before claiming a general gain.
Written for commit 6cadab8. Summary will update on new commits.
Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.
This PR includes no changesets
When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types
This PR is from an external contributor and must be approved by a stagehand team member with write access before CI can run.
Approving the latest commit mirrors it into an internal PR owned by the approver.
If new commits are pushed later, the internal PR stays open but is marked stale until someone approves the latest external commit and refreshes it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The existing implementation tries a lexical top-30 shortlist before asking Jev about the full candidate view. On lists with more than 40 candidates, repeated generic terms can fill those 30 slots and exclude a rare, relevant row.
This change keeps every candidate selected by the existing scorer. When all 30 slots are filled, it adds BM25's best candidate only if that candidate is outside the current shortlist. The list therefore contains at most 31 elements. If fewer than 30 candidates have lexical overlap, the current shortlist is returned without running BM25. Existing role views, tokenization, field weights, exact-name handling, and full-list retry remain in place. This may help when a rare, relevant term is missing from the existing shortlist while preserving its current candidates.
Measurements
Real websites
Stagehand captured accessibility snapshots in a local browser. Sites can change between runs, so counts below describe one capture.
For 30 Hacker News "hide" links labeled by author and 20 Books to Scrape basket buttons labeled by title, the correct target was in the top 30 for 50/50 with the current scorer, 50/50 with BM25, and 50/50 with the combined list. BM25 reordered the first Hacker News comments link, but both scorers already included it. Jev selected the intended GitHub New issue and Sapiens basket controls from real snapshots; no controls were clicked.
The real website sample shows equal recall and provides no evidence of an accuracy gain. The Acme improvement and first-record regression come from deliberately constructed cases. One added option can still change Jev's answer on an untested page. A larger labeled set with complete act results is needed before claiming a general success-rate or latency improvement.
Verification
extensionpack --checkpassed.Summary by cubic
Supplements Jev's lexical top-30 shortlist with BM25's best candidate so a rare, relevant row isn't excluded when repeated generic terms fill all 30 slots.
scoreCandidatesBm25andshortlistWithBm25inbm25.ts; the scorer has no new dependency.flattenTextandtokensfromtree.ts;pick.tsnow uses the new helper, and existing role views, tokenization, field weights, exact-name handling, and full-list retry are untouched.Measurements
Written for commit 6cadab8. Summary will update on new commits.