Skip to content

fix(review): bound scheduled PR review intake across repositories - #659

Merged
neubig merged 3 commits into
fix/653-scheduled-review-retryfrom
openhands/issue-656
Sep 25, 2026
Merged

neubig merged 3 commits into
fix/653-scheduled-review-retryfrom
openhands/issue-656

Conversation

@all-hands-bot

@all-hands-bot all-hands-bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor
  • A human has tested these changes.

Why

The scheduled reviewer-request scan from #654 drains every outstanding all-hands-bot request it finds, so one scan can start a conversation for every eligible PR. The four main OSS repositories currently have hundreds of open non-draft PRs lacking a review on their current head. Requesting all of them at once overloaded the 15 GiB OSS Agent Canvas VM: four simultaneous agent jobs plus retained Docker runtimes caused an OOM restart on 2026-09-22.

Summary

  • Bound a scheduled scan to a small configurable maximum of new review conversations, max_new_per_run, defaulting to 2 and counted across every configured repository rather than reset per repository. Eligible candidates are collected with the reviewer-request event and head SHA fix(review): resume requested reviews on a scheduled check scan #654 already keys deliveries on, ordered by the oldest outstanding request and then by repository and pull-request number, and started through the existing dispatcher.deliver. The bound counts conversations a scan starts, so a delivery that only deduplicates or reports an already-running conversation consumes no slot, and a candidate whose exact-head checks are pending or failing posts its gate disposition and consumes none either.
  • Keep the rest of the scan intact once the maximum is reached: the remaining candidates are still evaluated for exact-head check dispositions and gate comments, completed reviews are still reconciled, and maintainer handoff is unchanged. A dispatch that raises is reported without aborting the scan or the other repositories. The explicit review_requested event path and the trigger-label scan keep their current unbounded behavior.
  • Wire the bound through the worker's rendered config with the max_new_per_run key the other automations already use, document it in SKILL.md / README.md, and bump the catalog entry and bundle 1.6.0 → 1.7.0 with regenerated catalogs.

Details:

  • skills/github-pr-reviewer/scripts/worker.py
    • ReviewIntake holds the per-scan budget and a queue of eligible candidates. drain() sorts by (created_at, repository, number) and starts at most max_new_per_run of them, reusing dispatcher.deliver and its keyed dedupe. A dispatch that raises is collected and reported after the drain so later candidates still get their turn.
    • run_scan(dispatcher) is the shipped entrypoint: it creates one shared ReviewIntake, runs every configured repository, then drains once. The drain runs even when a repository raises, and the first failure is what the scan reports.
    • PullRequestReviewer.intake returns the shared scan intake when run_scan set one, otherwise a private lazily-created intake, so a standalone run() still bounds its own scan. In scheduled mode a green candidate is registered with the intake; in event mode it dispatches immediately and is never bounded.
  • skills/github-pr-reviewer/scripts/main.py — MAX_NEW_PER_RUN = 2 default plus max_new_per_run config parsing, rejecting a boolean (a Python bool is an int) and any value below 1.
  • automations/catalog/github-pr-reviewer/manifest.json — maxNewPerRun setup field (min 1, default 2), version bump, description/example/message update; automations/bundle-index.js, skills/index.js, and tests/fixtures/automations/github-pr-reviewer.json regenerated with npm run build.

Stacking

Base is fix/653-scheduled-review-retry, the head branch of #654, so GitHub records this as a stacked PR and it will retarget to main when #654 merges. #651 and #654 must merge first. This PR resumes the branch openhands/issue-656 after the earlier resolver pushed it but timed out before opening a PR; the unrelated integrations/catalog/granola.json and integrations/catalog-index.js changes that had been picked up from main were removed, so the diff below is only the bounded intake and its tests/docs.

Issue Number

Part of #656

How to Test

  1. uv run pytest -q tests/test_github_reviewer_delivery.py — expect 62 passed.
  2. uv run pytest -q — expect 1032 passed, 23 skipped (the base fbfece9 runs 1020 passed, 23 skipped; the difference is this PR's new cases).
  3. uv run python scripts/sync_extensions.py --check — expect clean (the pre-existing, non-blocking issue-duplicate-checker coverage warning is unrelated).
  4. npm run build — expect no drift in automations/bundle-index.js or skills/index.js.

Live Canvas / live-API validation

The Canvas automation service on this host was not reachable for a create/dispatch mutation (127.0.0.1:8001 returned 403 for the available key), so there is no live Canvas scheduled run to report. A read-only one-off harness instead loaded the shipped bundle from automations/catalog/github-pr-reviewer/manifest.json and drove the real PullRequestReviewer.run() scheduled path plus the real shared-intake drain against the live GitHub API: the dispatcher recorded what would have been delivered and the gate comment upsert was recorded instead of written, so no agent started and no comment was posted. That harness was evidence for #656, not part of the shipped automation, so it is not in this PR; the live result it produced is recorded below. The harness itself remains available on factory/reviewer-continuous-canary.

"open_pull_requests_inspected": {
  "OpenHands/extensions": 83,
  "OpenHands/OpenHands": 429,
  "OpenHands/software-agent-sdk": 262,
  "OpenHands/automation": 48
},
"registered_green_candidates": 1,
"drain_order": [
  { "repository": "OpenHands/OpenHands", "pr": 16890,
    "requested_at": "2026-09-23T02:10:14Z" }
],
"conversations_started_count": 1,
"gate_dispositions": [
  { "repository": "OpenHands/software-agent-sdk", "pr": 5058,
    "marker": "<!-- openhands-review-gate:blocked:adfeb8fd... -->" }
]

The live scan is honest about the gate, so only one candidate was green at the time of that run and the bound was not the limiting factor on live data. A second pass over the same live PRs, request events, head SHAs, and repositories forced each outstanding request eligible so the real drain is shown stopping at the bound on real data (only the gate flag is overridden; ordering, delivery keys, and the drain are the shipped code):

max_new_per_run=2 -> 2 conversations started (oldest first: sdk#5058 2026-09-22T12:24:25Z, then OpenHands#16890 2026-09-23T02:10:14Z)
max_new_per_run=1 -> 1 conversation started  (sdk#5058 only)
max_new_per_run=3 -> 2 conversations started (only 2 candidates existed)

That is the bounded intake: one scan starts no more than the configured number of conversations, in oldest-request order, and a later scan reaches the remainder. The remaining acceptance evidence — a review requested while a check is pending and completed after it goes green on a live deployment — requires the oss-agent-canvas automation to switch to the cron trigger, which is post-live-test deployment work and out of scope here.

Video/Screenshots

Not applicable: this changes automation worker dispatch logic, not a GUI.

Notes

  • The bound is a scheduling bound only. It does not create a queue or runtime layer: candidates are collected in memory for the duration of one scan and drained through the existing delivery machinery.
  • A candidate whose dispatch raises is reported in a single RuntimeError after the drain, so a flaky agent server does not silently drop the rest of the backlog.
  • The catalog default of 2 can be raised per deployment once the VM can hold more simultaneous agents.

Generated by OpenHands AI on behalf of the user.

@neubig

neubig commented Sep 23, 2026

Copy link
Copy Markdown
Member

@all-hands-bot #659 now conflicts with its updated base #654 (head fbfece9). Please restack this focused bounded-intake PR onto that head, preserve its behavior, run full CI, and push. #660 is stacked on #659, so keep its dependent commits out of this PR and do not merge. The clean combined canary branch in OpenHands/extensions at factory/reviewer-continuous-canary may help compare the intended result.

@neubig

neubig commented Sep 23, 2026

Copy link
Copy Markdown
Member

@all-hands-bot The restack is mergeable and CI green. Please remove .pr/canary_bounded_intake.py from this PR: it is a 197-line one-off evidence harness, not part of the shipped automation, and the PR should stay focused on the per-scan bound. Keep the live result in the PR description/comment, rerun checks, and do not merge.

@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Restacked #659 onto #654's head fbfece9 and pushed. New head: b55d263.

What the restack contains

openhands/issue-656 now sits directly on fbfece9 with exactly its own two commits and nothing from #660:

b55d263 chore(review): add read-only bounded-intake canary for #656
5c5a582 fix(review): bound scheduled PR review intake across repositories
fbfece9 (base, #654 head)

The old head's first two commits (1be2b44, 014659e) were the same logical changes as #654's own 28f025f/fbfece9, but on a pre-merge parent. They are dropped, so the diff is only the bounded intake and its canary — 10 files changed, 841 insertions(+), 51 deletions(-), the same shape as before the restack. #660's f573d0a/c01bac7 are not present, and git log on the branch shows no #660 commit.

One conflict resolution, and it is the one the canary branch implies

Three files conflicted: automations/bundle-index.js, skills/index.js, and skills/github-pr-reviewer/SKILL.md. The first two are generated, so I resolved them by running npm run build against the new base rather than hand-merging. SKILL.md was a both-added region: I kept the new base's check or workflow run wording and appended only the intake's own max_new_per_run troubleshooting row.

The one substantive adjustment was in tests/test_github_reviewer_delivery.py. The new base's gate reads workflow runs as well as check runs, so _scan_reviewers needs the run.workflow_runs = lambda sha: [] stub that the intake commit predates. Without it, test_the_explicit_request_event_path_is_not_bounded fails with StopIteration because the reviewer reaches for the un-stubbed workflow_runs during the gate. This is exactly what factory/reviewer-continuous-canary does — its _scan_reviewers has the same stub — and I confirmed the intake commit there (a78e3dc, parent fbfece9) is what this branch now matches: worker.py and main.py are byte-identical to it, and my intake commit's tree matches a78e3dc's exactly. I did not bring over any of that branch's #660 work.

Behavior preserved

Same bound, same semantics as before: at most max_new_per_run (default 2) new conversations per scan, counted across all repositories, ordered by oldest outstanding request then repository and PR number, delivered through the existing dispatcher.deliver. Reaching the bound still evaluates the remaining candidates for gate dispositions, and the explicit review_requested path and trigger-label scan stay unbounded. The catalog version bump 1.6.0 → 1.7.0 and the regenerated catalogs carry over.

CI

  • uv run --group test pytest tests/ — 1032 passed, 23 skipped.
  • uv run --group test pytest tests/test_github_reviewer_delivery.py — 62 passed.
  • npm run build — no drift; git status clean after regeneration.
  • uv run python scripts/sync_extensions.py --check — clean apart from the pre-existing, non-blocking issue-duplicate-checker coverage warning.
  • Read-only live canary (.pr/canary_bounded_intake.py) against the live API: the real drain stops at the bound, oldest-first — max_new_per_run=2 starts sdk#5058 then OpenHands#17598 and stops; 3 starts all three. No agent started and no comment posted.

GitHub CI on the new head: Validate PR description, Deprecation deadlines, pr-title lint/label all pass, and the PR is MERGEABLE. The Tests workflow does not appear on this PR or on #654; its pull_request: branches: ["*"] filter does not match a slashed base branch, which is pre-existing and unrelated to this change. I ran the suite locally instead.

I did not merge, and did not touch #660.

Generated by OpenHands AI on behalf of the user.

@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Done — the canary harness is out of the PR and the checks are green.

Change: pushed 7609cf1 (chore(review): drop the one-off bounded-intake canary harness) to openhands/issue-656, which is the only commit on top of the previous head b55d263 and deletes .pr/canary_bounded_intake.py (197 lines). No shipped-automation file changed in that commit. The PR diff is now 9 files, all of them the per-scan bound and its docs/tests/fixtures:

  • skills/github-pr-reviewer/scripts/worker.py
  • skills/github-pr-reviewer/scripts/main.py
  • skills/github-pr-reviewer/SKILL.md, README.md
  • automations/catalog/github-pr-reviewer/manifest.json
  • automations/bundle-index.js, skills/index.js
  • tests/fixtures/automations/github-pr-reviewer.json, tests/test_github_reviewer_delivery.py

Live result kept. The PR description still carries the recorded live-API evidence: the read-only scan over the four repositories (extensions 83, OpenHands 429, software-agent-sdk 262, automation 48 open PRs inspected, 1 green candidate, 1 conversation started, 1 gate disposition) and the forced-eligible drain demo (max_new_per_run=2 → 2, =1 → 1, =3 → 2, oldest-request first). I reworded the section so it no longer points at a file in this PR: it now describes the run as a one-off harness and notes the harness itself is preserved on factory/reviewer-continuous-canary, where it still lives alongside canary_unrequested_scan.py.

One correction while I was in there. The ## How to Test counts in the description were stale. I reran both suites on the current head and updated them: tests/test_github_reviewer_delivery.py is 62 passed, and the full suite is 1032 passed, 23 skipped (I also ran the base fbfece9 for comparison: 1020 passed, 23 skipped, so the difference is this PR's new cases). The old text said 56 / 1024.

Checks rerun on 7609cf1 (all pass, no pending or failed):

  • Validate PR description — pass
  • check (Deprecation deadlines) — pass
  • pr-title / Lint PR title (conventional) and pr-title / Apply conventional-type label — pass

Locally I also confirmed uv run pytest -q tests/test_github_reviewer_delivery.py (62 passed), uv run pytest -q (1032 passed, 23 skipped), uv run python scripts/sync_extensions.py --check (clean apart from the pre-existing, non-blocking issue-duplicate-checker coverage warning) and npm run build (no drift in automations/bundle-index.js or skills/index.js). The PR is MERGEABLE with mergeStateStatus: CLEAN.

Not merging, as asked. Two things for your awareness, neither of which I changed:

  1. feat(review): review unrequested green PRs on the scheduled scan #660 (openhands/issue-658) now shows as CONFLICTING against its base. That is not new and not caused by this commit: its head is based on 23ee3e7, which predates the restack onto fbfece9, and it conflicts in the shared files both PRs touch (bundle-index.js, manifest.json, worker.py, skills/index.js, the fixtures, and both test files). I checked the conflict set against the old head b55d263 and the new head 7609cf1 and it is identical, so this commit neither adds nor removes conflicts. It will need its own restack onto the new openhands/issue-656 head when you want it.
  2. feat(review): review unrequested green PRs on the scheduled scan #660 still carries its own .pr/canary_bounded_intake.py plus .pr/canary_unrequested_scan.py. I left both alone since you scoped this to fix(review): bound scheduled PR review intake across repositories #659, but the same reasoning you gave here would apply there if you want the stacked PR cleaned up too — say the word and I will do it.

Generated by OpenHands AI on behalf of the user.

@neubig

neubig commented Sep 23, 2026

Copy link
Copy Markdown
Member

@all-hands-bot After extensions #654 is updated, update this stacked PR onto the current head of its base branch fix/653-scheduled-review-retry, resolve conflicts while preserving the bounded cross-repository intake behavior, run validation, and push. Do not merge.

@all-hands-bot
all-hands-bot force-pushed the fix/653-scheduled-review-retry branch from fbfece9 to d63d9d7 Compare September 23, 2026 19:43
all-hands-bot and others added 3 commits September 23, 2026 19:48
A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)
Loads the shipped catalog bundle and drives the real scheduled scan and
shared-intake drain against the live GitHub API. It starts no agent and posts
no comment: the dispatcher and the gate comment upsert record what would have
happened, so the canary is read-only. Development-only, kept in .pr/.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit 9396c24)
The 197-line read-only canary was evidence for #656, not part of the shipped
automation, so it does not belong in this focused per-scan bound PR. The live
result stays recorded in the PR description; the harness itself is preserved on
factory/reviewer-continuous-canary.

Co-authored-by: openhands <openhands@all-hands.dev>
@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Rebased #659 onto the current head of fix/653-scheduled-review-retry (d63d9d7) and pushed. New head: 50fb986.

What the base move changed

#654 was force-pushed while I was working: its head went fbfece9 → d63d9d7, which adds 9c62801 "gate scheduled discovery on required checks only" and rewrites the two commits above it. The base now reads the PR-scoped GraphQL isRequired set and gates only those checks, with a fallback to every current-head check and workflow run when the required set cannot be read. Its gate entry point also changed shape:

  • _gate_head(pr, scheduled) → _gate_head(pr, requested=False, scheduled=False), where requested=True is the intake-policy exception that bypasses the gate entirely for an explicit all-hands-bot request.
  • _classify_check_runs(sha) → _classify_check_runs(pr), split into _classify_required_runs / _classify_all_runs.

I rebased the branch's three commits onto d63d9d7, dropping #654's duplicated commits as before.

How the two behaviors were reconciled

worker.py auto-merged, and the result is correct — the required-check gate and the bounded intake compose cleanly:

  • The gate call is the base's new form, _gate_head(pr, requested=event_mode, scheduled=not event_mode), so an explicit request still bypasses the required-check gate and a scheduled scan still gates on it.
  • The intake then registers the scheduled candidate and dispatches the explicit one immediately, exactly as before: if event_mode: self._start_review(...) else: self.intake.register(record).

The ReviewIntake class and run_scan entrypoint are byte-identical to the pre-rebase versions, so #659's bounded cross-repository scheduled intake is unchanged: at most max_new_per_run (default 2) new conversations per scan, counted across all repositories, oldest outstanding request first, drain-through-dispatcher.deliver, unstarted deliveries and gate dispositions consuming no slot, and the explicit review_requested path never bounded.

Three files conflicted:

  • skills/github-pr-reviewer/SKILL.md — a both-added troubleshooting block. I kept the base's three required-check rows (Review paused with a failing-check comment, Review reported waiting on checks, Optional workflow failed but no review was paused) and appended only the intake's own max_new_per_run row.
  • automations/bundle-index.js and skills/index.js — generated. I resolved them by taking the base version and re-running npm run build, rather than hand-merging, so they contain the base's required-check gate text plus the intake bound. npm run build is clean with no drift.

The docs now describe both behaviors: the base's "Gates each scheduled review on the current head's GitHub-required checks … an explicit all-hands-bot review request bypasses the gate by design" sits alongside the intake's "Bounded intake per scheduled scan" section.

Validation

  • uv run --group test pytest tests/ — 1042 passed, 23 skipped.
  • uv run --group test pytest tests/test_github_reviewer_delivery.py — 69 passed.
  • Targeted run of the gate + intake tests (required, not_bounded, max_new, one_scan, explicit_request_bypasses, scheduled_scan) — 20 passed, so the required-check gate and the bound hold together.
  • npm run build — no drift; uv run python scripts/sync_extensions.py --check — clean apart from the pre-existing, non-blocking issue-duplicate-checker coverage warning.
  • GitHub CI on 50fb986: Validate PR description, Deprecation deadlines, pr-title lint/label all pass, and the PR is MERGEABLE.

git log on the branch shows only the three #659 commits and no #660 content. I did not reintroduce the .pr/ canary harness that a later commit intentionally dropped. I did not merge and did not touch #660.

Generated by OpenHands AI on behalf of the user.

all-hands-bot pushed a commit that referenced this pull request Sep 23, 2026
Restacked onto the current #659 head (openhands/issue-656). This keeps the
continuous unrequested green-PR discovery and the corrected base behavior it
builds on: scheduled candidates gate on GitHub-required checks (falling back to
every current-head check and workflow run when the required signal is
unavailable), and an explicit all-hands-bot request bypasses the CI gate.

The scheduled scan examined every open, non-draft PR. Classifying each
unrequested head costs a review read plus the exact-head check and workflow
reads, so one run read one list per PR across the largest repository and posted
a managed gate comment for every red or pending head, which the live canary of
#660 exposed as a blocking scalability bug.

- The unrequested part of a scheduled scan is now a bounded, rotating window:
  at most SCAN_WINDOW (10) unrequested PRs per repository, starting where the
  previous scan stopped. The position is retained in the existing Automation KV
  store under a per-repository review-scan:{owner}__{repo} key, so successive
  scans rotate through the whole backlog instead of reading one pull request per
  open PR. With no KV store the position is kept in memory.
- Explicit all-hands-bot review requests and trigger labels are never subject to
  the window: every explicit candidate is examined on every scan, and an
  explicit request still bypasses the CI gate.
- A gate stop on an unrequested head posts no managed comment, so a scan over a
  large backlog cannot storm the PRs with comments. The managed comment stays for
  explicit requests and labeled heads, which still get their blocked or waiting
  explanation.
- Remove the one-off .pr canary script and bump the catalog entry and bundle
  1.7.0 -> 1.8.0.

Tests cover rotation and KV persistence, explicit priority over the window,
the absence of unrequested gate comments, and the bounded pull-request reads,
alongside the base required-check and explicit-request-bypass tests.
@neubig
neubig merged commit 0f3a2df into fix/653-scheduled-review-retry Sep 25, 2026
6 checks passed
@neubig
neubig deleted the openhands/issue-656 branch September 25, 2026 20:47
neubig pushed a commit that referenced this pull request Sep 25, 2026
Restacked onto the current #659 head (openhands/issue-656). This keeps the
continuous unrequested green-PR discovery and the corrected base behavior it
builds on: scheduled candidates gate on GitHub-required checks (falling back to
every current-head check and workflow run when the required signal is
unavailable), and an explicit all-hands-bot request bypasses the CI gate.

The scheduled scan examined every open, non-draft PR. Classifying each
unrequested head costs a review read plus the exact-head check and workflow
reads, so one run read one list per PR across the largest repository and posted
a managed gate comment for every red or pending head, which the live canary of
#660 exposed as a blocking scalability bug.

- The unrequested part of a scheduled scan is now a bounded, rotating window:
  at most SCAN_WINDOW (10) unrequested PRs per repository, starting where the
  previous scan stopped. The position is retained in the existing Automation KV
  store under a per-repository review-scan:{owner}__{repo} key, so successive
  scans rotate through the whole backlog instead of reading one pull request per
  open PR. With no KV store the position is kept in memory.
- Explicit all-hands-bot review requests and trigger labels are never subject to
  the window: every explicit candidate is examined on every scan, and an
  explicit request still bypasses the CI gate.
- A gate stop on an unrequested head posts no managed comment, so a scan over a
  large backlog cannot storm the PRs with comments. The managed comment stays for
  explicit requests and labeled heads, which still get their blocked or waiting
  explanation.
- Remove the one-off .pr canary script and bump the catalog entry and bundle
  1.7.0 -> 1.8.0.

Tests cover rotation and KV persistence, explicit priority over the window,
the absence of unrequested gate comments, and the bounded pull-request reads,
alongside the base required-check and explicit-request-bypass tests.
neubig pushed a commit that referenced this pull request Sep 25, 2026
* fix(review): bound scheduled PR review intake across repositories

A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)

* chore(review): add read-only bounded-intake canary for #656

Loads the shipped catalog bundle and drives the real scheduled scan and
shared-intake drain against the live GitHub API. It starts no agent and posts
no comment: the dispatcher and the gate comment upsert record what would have
happened, so the canary is read-only. Development-only, kept in .pr/.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit 9396c24)

* chore(review): drop the one-off bounded-intake canary harness

The 197-line read-only canary was evidence for #656, not part of the shipped
automation, so it does not belong in this focused per-scan bound PR. The live
result stays recorded in the PR description; the harness itself is preserved on
factory/reviewer-continuous-canary.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
neubig pushed a commit that referenced this pull request Sep 25, 2026
* fix(review): resume requested reviews on a scheduled check scan

The OSS reviewer's only trigger was the GitHub review_requested event. #651
stopped dispatching an agent while the exact head's checks were pending or
failing, and its waiting comment promised a scheduled retry that deployment did
not have, so a request arriving during CI could be left unreviewed.

Reuse the existing worker and one automation record. In scheduled mode the scan
now also considers every open, non-draft PR that still holds an outstanding
all-hands-bot review request, keyed by that request's own review_requested
event so repeated scans reuse one conversation and one review. The explicit
request event path is unchanged for event-only deployments.

The waiting and blocked comments now name the retry the deployment actually
has: a scheduled scan where a cron trigger exists, and another review request
where only the event trigger does.

Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and
adds scheduled-scan, duplicate-scan, draft, and gate-message tests.

Part of #653. Stacked on #651.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(review): reword the managed gate comment across a trigger change

A pending PR first gated by an event-only deployment kept its 'request
all-hands-bot again' comment after the automation was switched to a cron scan,
because the gate returned early on a matching state:sha marker without comparing
the body. The comment therefore named a retry that deployment no longer had.

The gate now compares the full managed body for the same marker and rewrites the
comment in place when it changed, so an event-only comment becomes the scheduled
one once the scan exists and an identical body stays a no-op. No second comment
is created.

The event-only wording also makes the retry explicit: the outstanding request
must be removed and re-requested, since GitHub will not accept a second request
for a reviewer who is already requested.

Adds a focused event-then-cron regression plus a stale-body rewrite test, and
documents the in-place reword in SKILL.md and README.md.

Part of #653.

Co-authored-by: openhands <openhands@all-hands.dev>

* fix(review): bound scheduled PR review intake across repositories (#659)

* fix(review): bound scheduled PR review intake across repositories

A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)

* chore(review): add read-only bounded-intake canary for #656

Loads the shipped catalog bundle and drives the real scheduled scan and
shared-intake drain against the live GitHub API. It starts no agent and posts
no comment: the dispatcher and the gate comment upsert record what would have
happened, so the canary is read-only. Development-only, kept in .pr/.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit 9396c24)

* chore(review): drop the one-off bounded-intake canary harness

The 197-line read-only canary was evidence for #656, not part of the shipped
automation, so it does not belong in this focused per-scan bound PR. The live
result stays recorded in the PR description; the harness itself is preserved on
factory/reviewer-continuous-canary.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com>
neubig pushed a commit that referenced this pull request Sep 25, 2026
Restacked onto the current #659 head (openhands/issue-656). This keeps the
continuous unrequested green-PR discovery and the corrected base behavior it
builds on: scheduled candidates gate on GitHub-required checks (falling back to
every current-head check and workflow run when the required signal is
unavailable), and an explicit all-hands-bot request bypasses the CI gate.

The scheduled scan examined every open, non-draft PR. Classifying each
unrequested head costs a review read plus the exact-head check and workflow
reads, so one run read one list per PR across the largest repository and posted
a managed gate comment for every red or pending head, which the live canary of

- The unrequested part of a scheduled scan is now a bounded, rotating window:
  at most SCAN_WINDOW (10) unrequested PRs per repository, starting where the
  previous scan stopped. The position is retained in the existing Automation KV
  store under a per-repository review-scan:{owner}__{repo} key, so successive
  scans rotate through the whole backlog instead of reading one pull request per
  open PR. With no KV store the position is kept in memory.
- Explicit all-hands-bot review requests and trigger labels are never subject to
  the window: every explicit candidate is examined on every scan, and an
  explicit request still bypasses the CI gate.
- A gate stop on an unrequested head posts no managed comment, so a scan over a
  large backlog cannot storm the PRs with comments. The managed comment stays for
  explicit requests and labeled heads, which still get their blocked or waiting
  explanation.
- Remove the one-off .pr canary script and bump the catalog entry and bundle
  1.7.0 -> 1.8.0.

Tests cover rotation and KV persistence, explicit priority over the window,
the absence of unrequested gate comments, and the bounded pull-request reads,
alongside the base required-check and explicit-request-bypass tests.
neubig pushed a commit that referenced this pull request Sep 25, 2026
fix(review): bound the unrequested scan to a rotating KV-backed window

Restacked onto the current #659 head (openhands/issue-656). This keeps the
continuous unrequested green-PR discovery and the corrected base behavior it
builds on: scheduled candidates gate on GitHub-required checks (falling back to
every current-head check and workflow run when the required signal is
unavailable), and an explicit all-hands-bot request bypasses the CI gate.

The scheduled scan examined every open, non-draft PR. Classifying each
unrequested head costs a review read plus the exact-head check and workflow
reads, so one run read one list per PR across the largest repository and posted
a managed gate comment for every red or pending head, which the live canary of

- The unrequested part of a scheduled scan is now a bounded, rotating window:
  at most SCAN_WINDOW (10) unrequested PRs per repository, starting where the
  previous scan stopped. The position is retained in the existing Automation KV
  store under a per-repository review-scan:{owner}__{repo} key, so successive
  scans rotate through the whole backlog instead of reading one pull request per
  open PR. With no KV store the position is kept in memory.
- Explicit all-hands-bot review requests and trigger labels are never subject to
  the window: every explicit candidate is examined on every scan, and an
  explicit request still bypasses the CI gate.
- A gate stop on an unrequested head posts no managed comment, so a scan over a
  large backlog cannot storm the PRs with comments. The managed comment stays for
  explicit requests and labeled heads, which still get their blocked or waiting
  explanation.
- Remove the one-off .pr canary script and bump the catalog entry and bundle
  1.7.0 -> 1.8.0.

Tests cover rotation and KV persistence, explicit priority over the window,
the absence of unrequested gate comments, and the bounded pull-request reads,
alongside the base required-check and explicit-request-bypass tests.

Co-authored-by: openhands <openhands@all-hands.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type: fix A bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants