Skip to content

fix(review): resume requested reviews on a scheduled check scan - #654

Merged
neubig merged 3 commits into
mainfrom
fix/653-scheduled-review-retry
Sep 25, 2026
Merged

neubig merged 3 commits into
mainfrom
fix/653-scheduled-review-retry

Conversation

@all-hands-bot

@all-hands-bot all-hands-bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor
  • A human has tested these changes.

Why

The live OSS reviewer's only trigger was the GitHub review_requested event. #651 stopped dispatching an agent while the exact head's checks were pending or failing, and its waiting comment promised a scheduled retry that this deployment did not have. A request that arrived during CI could therefore be left unreviewed unless a human removed and re-requested all-hands-bot, which undermines automatic coverage of non-draft PRs.

Summary

  • Reuse the existing skills/github-pr-reviewer worker and one reviewer automation record: in scheduled mode the scan now also considers every open, non-draft PR that still holds an outstanding all-hands-bot review request, keyed by that request's own review_requested event so repeated scans reuse one conversation and one review. The explicit-request event path is unchanged and stays usable in event-only deployments. No second automation record, no new GitHub webhook types, no check_run parsing, and no deployment-specific code in this repo.
  • Name the retry the deployment actually has: a scheduled run says the next scan retries; an event-only run asks for another all-hands-bot request instead of promising a scan. Bump the catalog entry and bundle 1.5.0 → 1.6.0, move the default catalog schedule to */5 * * * *, and document the retry contract in SKILL.md / README.md.
  • Add tests for pending-to-green across two scans, duplicate scans reusing one delivery key, a draft-with-request, an unlabeled+unrequested PR, the unchanged label path, and both event-path gate comments refusing to promise a scan.

Details:

  • skills/github-pr-reviewer/scripts/worker.py
    • _outstanding_review_request(pr) reads requested_reviewers off the list endpoint the scheduled scan already fetches. It is the live set, so an answered or withdrawn request simply drops out; open, non-draft PRs only.
    • In scheduled mode, a PR with an outstanding request but no trigger label is keyed by its own review_requested event, so repeated scans produce the same delivery ({event_id}:{sha}) and the existing keyed subject/delivery dedupe reuses one conversation and one review.
    • _gate_body() takes a scheduled flag and names the retry that actually exists.
    • The eligibility gate, latest-run-per-logical-check policy, keyed dedupe, completed-review short-circuit, verdict parsing, maintainer handoff, and secret handling are untouched.
  • automations/catalog/github-pr-reviewer/manifest.json — version bump, description/example update, default schedule */5 * * * *; automations/bundle-index.js and skills/index.js regenerated with npm run build.

Stacking

Base is fix/644-reject-failed-checks, the head branch of #651, so GitHub records this as a stacked PR and it will retarget to main when #651 merges. #651 must merge first. I will update the base when #651's check-ordering fix lands.

Issue Number

Part of #653

How to Test

Required. Share the steps for the reviewer to be able to test your PR.

  1. uv run pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py tests/test_automation_setup.py tests/test_catalogs.py tests/test_interface_manifest.py — expect 209 passed, 17 skipped.
  2. uv run pytest -q — expect 1010 passed, 23 skipped.
  3. uv run python scripts/sync_extensions.py --check — expect clean (the pre-existing, non-blocking issue-duplicate-checker coverage warning is unrelated).
  4. npm run build — expect no drift in automations/bundle-index.js or skills/index.js.

Live Canvas / live-API validation

The Canvas automation service on this host was not reachable for a create/dispatch mutation (127.0.0.1:8001 returned 403 for the available key, and ${OPENHANDS_URL} resolves to an ERR_NGORK_3200 tunnel), so there is no live Canvas run to report. Instead I loaded the shipped bundle from automations/catalog/github-pr-reviewer/manifest.json and drove the real PullRequestReviewer scan against the live GitHub API with a read-only reviewer (no dispatch, no comment posted) across all four configured repos:

=== OpenHands/extensions ===
  open PRs: 83
  outstanding all-hands-bot requests: 0

=== OpenHands/OpenHands ===
  open PRs: 434
  outstanding all-hands-bot requests: 1
    #17584 draft=False head=abbbf032 state=green names=[]

=== OpenHands/software-agent-sdk ===
  open PRs: 256
  outstanding all-hands-bot requests: 1
    #5058 draft=False head=adfeb8fd state=blocked names=['Check OpenAPI Schema', 'agent-server-tests', 'pre-commit', 'sdk-tests']

=== OpenHands/automation ===
  open PRs: 48
  outstanding all-hands-bot requests: 1
    #374 draft=False head=27175d12 state=green names=[]

This exercises the new scan against live data: requested_reviewers is populated on real PRs, drafts are excluded, and each outstanding request resolves to a real exact head and gate disposition. The remaining acceptance evidence — a review requested while a check is pending and completed after it goes green on a live deployment — requires the oss-agent-canvas automation to switch to the cron trigger, which is post-live-test deployment work and not part of this change.

Video/Screenshots

Not applicable: this changes automation worker dispatch logic, not a GUI.

Notes

  • The five-minute interval is the deployment minimum shown in the catalog fixtures; the entry does not enforce it, the service does.
  • Review verdicts on real bot reviews may not always end with the parseable verdict line, so like fix(review): gate agent launch on exact-head checks #651 this scan can occasionally re-dispatch a stable, keyed conversation. Spending one slot is preferred over silently skipping an outstanding request.

Generated by OpenHands AI on behalf of the user.

@github-actions github-actions Bot added the type: fix A bug fix label Sep 22, 2026
@neubig

neubig commented Sep 22, 2026

Copy link
Copy Markdown
Member

One live-transition gap: _gate_comment() returns early when it finds the same state:sha marker. A pending PR first gated by the event-only deployment keeps the old “request again” comment after this automation is switched to cron, because the scheduled scan never updates the wording. The acceptance criteria call for the comment to name the retry condition actually deployed. Please update the managed comment when the body/retry mode changes (without creating a second comment), with a focused event-then-cron regression. Also make the event-only instruction explicit that an outstanding request must be removed and re-requested, since GitHub cannot request the same reviewer twice while it remains outstanding. No merge.

@neubig

neubig commented Sep 22, 2026

Copy link
Copy Markdown
Member

@all-hands-bot please address the event-to-cron managed-comment transition and the event-only retry wording described in my latest comment. Add focused regression coverage, keep the scope narrow, run relevant tests, and push. Do not merge. Live Canvas validation will follow.

@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Addressed both points and pushed to fix/653-scheduled-review-retry (commit d242ee4).

1. Event-to-cron managed-comment transition

_gate_comment() no longer returns early on a matching state:sha marker. When a managed marker matches, it now picks that comment and compares its full body against the body about to be posted:

  • identical body → no-op, no API call (still no duplicate comment);
  • changed body → PATCH /issues/comments/{id} in place, so the same head's comment now names the retry the currently deployed trigger actually has.

That covers the live gap: a pending PR first gated by the event-only deployment holds the "remove and re-request" comment on waiting:head, and the first scheduled scan rewrites it to "The scheduled scan retries" without creating a second comment. The forged-marker and unmarked-workflow-comment protections are unchanged — only comments this account authored with the gate marker are eligible.

2. Event-only retry wording is now explicit

Both event-path bodies now say the outstanding request must be removed and re-requested, since GitHub rejects a second request for a reviewer who is already requested:

  • blocked: "Fix the checks above and push. Then remove the outstanding all-hands-bot request and request all-hands-bot again: GitHub will not accept a second request while the first is still outstanding. The review starts on the updated head."
  • waiting: same remove-then-re-request instruction once every check reports a conclusion.

Regression coverage

  • test_reviewer_rewords_the_managed_comment_when_an_event_deployment_becomes_scheduled — event run gates the head and posts the event-only body, then a scheduled scan over the same head PATCHes that comment to the scheduled wording, with zero POSTs.
  • test_reviewer_rewrites_a_stale_gate_comment_body_for_the_same_head — same-head body drift is rewritten in place.
  • test_reviewer_does_not_duplicate_the_gate_comment_for_the_same_head — an identical body is still a no-op (asserts no POST and no PATCH).
  • Both event gate-comment tests now assert the remove-and-re-request phrasing and that no "scheduled" wording leaks into an event deployment.

Docs

SKILL.md and README.md state the in-place reword rule and the remove-then-re-request requirement; npm run build was rerun.

Verification

  • uv run pytest -q tests/test_github_reviewer_delivery.py — 44 passed.
  • uv run pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py tests/test_automation_setup.py tests/test_catalogs.py tests/test_interface_manifest.py tests/test_skills_catalog.py — 261 passed, 17 skipped.
  • uv run pytest -q — 1012 passed, 23 skipped.
  • uv run python scripts/sync_extensions.py --check — clean apart from the pre-existing non-blocking issue-duplicate-checker coverage warning.
  • npm run build — no drift beyond the regenerated bundle-index.js / skills/index.js.

Scope is limited to the gate-comment upsert, the two event-path bodies, their docs, and the focused tests; the scheduled-scan selection, dedupe keys, and check ordering are untouched. The catalog 1.6.0 bump from the earlier commit already covers this PR, so no additional version bump. Not merged, as requested.

Generated by OpenHands AI on behalf of the user.

@neubig

neubig commented Sep 22, 2026

Copy link
Copy Markdown
Member

@all-hands-bot Please rebase this stacked PR onto the current head of its base PR #651 before it is reviewed or deployed. The #654 branch still carries _check_run_order() using (started_at or "~{run_id}", run_id), while #651 now uses run ID as the primary key to fix both queued-rerun ordering directions. Preserve the narrow scheduled-retry change, regenerate the bundle/catalog, run focused and repository checks, and push. Do not merge. I am testing the current canary separately; please do not change its live deployment.

@all-hands-bot
all-hands-bot force-pushed the fix/653-scheduled-review-retry branch from d242ee4 to 014659e Compare September 22, 2026 21:37
@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Rebased and pushed to fix/653-scheduled-review-retry (force-with-lease, d242ee4 → 014659e). Not merged; the live canary deployment was not touched.

Rebase

git rebase --onto origin/fix/644-reject-failed-checks de8dd38 fix/653-scheduled-review-retry. The new base head 0c4b297 is now an ancestor of the branch, and the only diff against #651's head is this PR's own narrow scheduled-retry change (8 files, +374/−64).

The branch no longer carries the old ordering. _check_run_order() is now #651's run-ID-primary form:

      run_id = int(run.get("id") or 0)
      return (run_id, run.get("started_at") or "")

The (started_at or "~{run_id}", run_id) tuple and the ~{run_id:020d} fallback are gone from worker.py, from the two SKILL.md/README.md ordering paragraphs, and from both regenerated artifacts. Conflicts in tests/test_github_reviewer_delivery.py, automations/bundle-index.js, and skills/index.js were resolved by keeping #651's ordering content and reapplying this PR's scheduled-retry additions on top — the new test_reviewer_launches_when_a_success_supersedes_an_older_null_start_run and the scheduled-retry block now coexist in one file.

Scheduled-retry change preserved

_outstanding_review_request(), the scheduled-scan selection keyed by the request's own review_requested event, _gate_body(..., scheduled), and the event-to-cron in-place gate-comment reword are all intact. The catalog 1.6.0 manifest, */5 * * * * default schedule, and the README/SKILL retry contract are unchanged.

Verification

  • uv run --group test pytest -q tests/test_github_reviewer_delivery.py — 44 passed.
  • uv run --group test pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py tests/test_automation_setup.py tests/test_catalogs.py tests/test_interface_manifest.py tests/test_skills_catalog.py — 261 passed, 17 skipped.
  • uv run --group test pytest -q — 1012 passed, 23 skipped.
  • uv run --group test python scripts/sync_extensions.py --check — exit 0, only the pre-existing non-blocking issue-duplicate-checker coverage warning.
  • npm run build — regenerated automations/bundle-index.js and skills/index.js and committed them; a second npm run build produces no further diff, and the bundle now embeds return (run_id, run.get("started_at") or "") with no ~{run_id:020d}.

Generated by OpenHands AI on behalf of the user.

all-hands-bot added a commit that referenced this pull request Sep 23, 2026
A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)
@neubig

neubig commented Sep 23, 2026

Copy link
Copy Markdown
Member

@all-hands-bot #654 now conflicts with its base #651 after the exact-head workflow-run gate was fixed. Please restack this PR onto the current #651 head (6034e07), resolve only the resulting conflicts, regenerate generated artifacts if needed, and run CI. Keep the scheduled retry behavior focused. #659 is stacked on this branch, so preserve a clean history for its subsequent restack. Do not merge.

@all-hands-bot
all-hands-bot force-pushed the fix/653-scheduled-review-retry branch from 014659e to fbfece9 Compare September 23, 2026 04:02
@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Restacked onto the current #651 head 6034e07 and pushed to fix/653-scheduled-review-retry (force-with-lease, 014659e → fbfece9). Not merged.

Restack

git rebase --onto origin/fix/644-reject-failed-checks 0c4b297 fix/653-scheduled-review-retry. 6034e07 is now an ancestor of the branch, the history is linear (2 commits, no merge commits), and the diff against the new base head is this PR's own change only: 8 files, +374/−64 — identical in scope to the pre-restack diff.

Conflicts resolved (only the resulting ones)

  • skills/github-pr-reviewer/README.md, skills/github-pr-reviewer/SKILL.md — both sides edited the same feature bullet and the ordering paragraph. Kept fix(review): gate agent launch on exact-head checks #651's workflow-run text and the run-ID-primary ordering ("by the run ID ... with the start time as a tie-break"), then reapplied this PR's scheduled-scan bullet, the retry-naming paragraph, and the in-place reword paragraph on top.
  • automations/bundle-index.js, skills/index.js — generated; resolved by taking the base copy and re-running npm run build over the merged sources, per commit, so each commit is internally consistent.
  • worker.py and tests/test_github_reviewer_delivery.py auto-merged with no conflict.

The branch no longer carries the old ordering. _check_run_order() is #651's run-ID-primary form ((run_id, run.get("started_at") or "")), the ~{run_id:020d} fallback is gone from the sources, the docs, and both regenerated artifacts, and #651's new workflow_runs()/_latest_workflow_runs() current-head workflow gate is present. The scheduled-retry change is intact and still narrow: _outstanding_review_request(), the scheduled-scan selection keyed by the request's own review_requested event, _gate_body(..., scheduled), and the event-to-cron in-place gate-comment reword.

Verification

  • uv run --group test pytest -q tests/test_github_reviewer_delivery.py — 50 passed.
  • uv run --group test pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py tests/test_automation_setup.py tests/test_catalogs.py tests/test_interface_manifest.py tests/test_skills_catalog.py — 268 passed, 17 skipped.
  • uv run --group test pytest -q — 1020 passed, 23 skipped.
  • uv run --group test python scripts/sync_extensions.py --check — exit 0, only the pre-existing non-blocking issue-duplicate-checker coverage warning.
  • npm run build — no drift after the restack.

CI

gh pr checks 654 on head fbfece9 is green (Validate PR description, check, pr-title), and GitHub now reports the PR as MERGEABLE against fix/644-reject-failed-checks.

#659

#659 still targets this branch, so its own restack onto the new tip fbfece9 is needed; GitHub currently shows it CONFLICTING against the previous tip. I left that branch untouched so its history stays clean for that restack. The live canary deployment was not changed.

Generated by OpenHands AI on behalf of the user.

all-hands-bot added a commit that referenced this pull request Sep 23, 2026
A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)
@neubig

neubig commented Sep 23, 2026

Copy link
Copy Markdown
Member

@all-hands-bot Please update this stacked PR onto the current head of its base branch fix/644-reject-failed-checks (extensions #651), resolve conflicts while preserving this PR's scheduled retry behavior, run the focused/full validation, and push. Do not merge.

@all-hands-bot
all-hands-bot force-pushed the fix/653-scheduled-review-retry branch from fbfece9 to d63d9d7 Compare September 23, 2026 19:43
@all-hands-bot

Copy link
Copy Markdown
Contributor Author

Restacked onto the current #651 head 9c62801 and pushed to fix/653-scheduled-review-retry (force-with-lease, fbfece9 → d63d9d7). Not merged.

Restack

git rebase --onto origin/fix/644-reject-failed-checks 6034e07 fix/653-scheduled-review-retry. 9c62801 is now an ancestor of the branch, the history is linear (2 commits, no merge commits), and the diff against the new base is this PR's change only: 8 files, +355/−65.

Conflicts resolved

  • skills/github-pr-reviewer/SKILL.md, skills/github-pr-reviewer/scripts/worker.py — fix(review): gate agent launch on exact-head checks #651 rewrote the same gate section and changed _gate_head. Kept fix(review): gate agent launch on exact-head checks #651's required-check semantics verbatim and re-applied this PR's scheduled-scan bullet and retry wording on top. _gate_head(pr, requested=False, scheduled=False) now merges both: the explicit-request bypass from fix(review): gate agent launch on exact-head checks #651 (if requested: return "green", sha, no gate comment) and this PR's scheduled flag threaded into _gate_body, with run() calling _gate_head(pr, requested=event_mode, scheduled=not event_mode).
  • automations/bundle-index.js, skills/index.js — generated; resolved by re-running npm run build over the merged sources, per commit, so each commit is internally consistent.
  • README.md, manifest.json, fixture, and the rest of the test file auto-merged.

One semantic conflict, reconciled

#651 made an explicit all-hands-bot request bypass the CI gate entirely ("Dispatches an explicit all-hands-bot review request even when required CI is red or pending; the CI gate applies to scheduled discovery"), so an event-path gate comment is no longer reachable through run(). This PR's three event-path tests asserted that path, so they failed against the merged code. I reconciled them without weakening either behavior:

  • The two event-wording tests now assert the contract directly on _gate_body(..., scheduled=False) (it still names remove-and-re-request, still refuses to promise a scan).
  • test_reviewer_rewords_the_managed_comment_when_an_event_deployment_becomes_scheduled now seeds the comment an event-only deployment would have left (_gate_body(..., scheduled=False)), runs a scheduled scan, and asserts a single PATCH to the scheduled wording with zero POSTs — the same in-place transition the original test covered.

Preserved

#651's required-check gate (required_check_contexts, required-only classification, fallback to the full rollup, and the explicit-request bypass) is intact, and the old run-ID ordering (~{run_id:020d}) is gone from the sources and both artifacts. This PR's scheduled retry is intact and still narrow: _outstanding_review_request(), the scheduled-scan selection keyed by the request's own review_requested event, the */5 * * * * default, the catalog 1.6.0 bump, and the event-to-cron in-place gate-comment reword.

Validation

  • uv run --group test pytest -q tests/test_github_reviewer_delivery.py — 57 passed.
  • uv run --group test pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py tests/test_automation_setup.py tests/test_catalogs.py tests/test_interface_manifest.py tests/test_skills_catalog.py — 278 passed, 17 skipped.
  • uv run --group test pytest -q — 1030 passed, 23 skipped.
  • uv run --group test python scripts/sync_extensions.py --check — exit 0, only the pre-existing non-blocking issue-duplicate-checker coverage warning.
  • npm run build — no drift after the restack; both commits rebuild clean.

CI

gh pr checks 654 on head d63d9d7 is green (Validate PR description, check, pr-title), and GitHub reports the PR MERGEABLE against fix/644-reject-failed-checks.

#659 still targets this branch, so it needs its own restack onto the new tip; I left it untouched to keep its history clean. The live deployment was not changed.

Generated by OpenHands AI on behalf of the user.

all-hands-bot added a commit that referenced this pull request Sep 23, 2026
A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)
neubig pushed a commit that referenced this pull request Sep 25, 2026
* fix(review): bound scheduled PR review intake across repositories

A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)

* chore(review): add read-only bounded-intake canary for #656

Loads the shipped catalog bundle and drives the real scheduled scan and
shared-intake drain against the live GitHub API. It starts no agent and posts
no comment: the dispatcher and the gate comment upsert record what would have
happened, so the canary is read-only. Development-only, kept in .pr/.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit 9396c24)

* chore(review): drop the one-off bounded-intake canary harness

The 197-line read-only canary was evidence for #656, not part of the shipped
automation, so it does not belong in this focused per-scan bound PR. The live
result stays recorded in the PR description; the harness itself is preserved on
factory/reviewer-continuous-canary.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
@neubig
neubig deleted the branch main September 25, 2026 20:52
@neubig neubig closed this Sep 25, 2026
@neubig neubig reopened this Sep 25, 2026
@neubig
neubig changed the base branch from fix/644-reject-failed-checks to main September 25, 2026 20:52
openhands-agent and others added 3 commits September 25, 2026 20:53
The OSS reviewer's only trigger was the GitHub review_requested event. #651
stopped dispatching an agent while the exact head's checks were pending or
failing, and its waiting comment promised a scheduled retry that deployment did
not have, so a request arriving during CI could be left unreviewed.

Reuse the existing worker and one automation record. In scheduled mode the scan
now also considers every open, non-draft PR that still holds an outstanding
all-hands-bot review request, keyed by that request's own review_requested
event so repeated scans reuse one conversation and one review. The explicit
request event path is unchanged for event-only deployments.

The waiting and blocked comments now name the retry the deployment actually
has: a scheduled scan where a cron trigger exists, and another review request
where only the event trigger does.

Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and
adds scheduled-scan, duplicate-scan, draft, and gate-message tests.

Part of #653. Stacked on #651.

Co-authored-by: openhands <openhands@all-hands.dev>
A pending PR first gated by an event-only deployment kept its 'request
all-hands-bot again' comment after the automation was switched to a cron scan,
because the gate returned early on a matching state:sha marker without comparing
the body. The comment therefore named a retry that deployment no longer had.

The gate now compares the full managed body for the same marker and rewrites the
comment in place when it changed, so an event-only comment becomes the scheduled
one once the scan exists and an identical body stays a no-op. No second comment
is created.

The event-only wording also makes the retry explicit: the outstanding request
must be removed and re-requested, since GitHub will not accept a second request
for a reviewer who is already requested.

Adds a focused event-then-cron regression plus a stale-body rewrite test, and
documents the in-place reword in SKILL.md and README.md.

Part of #653.

Co-authored-by: openhands <openhands@all-hands.dev>
* fix(review): bound scheduled PR review intake across repositories

A scheduled scan resumed every outstanding all-hands-bot request it found and
started one conversation for each, so a first scan over the OSS backlog could
start an agent for hundreds of eligible pull requests at once and exhaust the
deployment. The scan now starts at most `max_new_per_run` new review
conversations, defaulting to 2, counted across every configured repository
rather than reset per repository.

Delivery reuses the existing machinery: eligible candidates are collected with
the request event and head SHA #654 already keys deliveries on, ordered by the
oldest outstanding request and then by repository and pull-request number, and
started through the same `dispatcher.deliver`. The bound counts conversations a
scan starts, so a deduplicated or already-running delivery reuses its runtime and
consumes no slot, and a candidate whose exact-head checks are pending or failing
posts its gate disposition and consumes none either. Reaching the bound still
evaluates the remaining candidates and their gate comments, and a dispatch that
raises is reported without aborting the scan. The explicit reviewer-request
event path and the trigger-label scan are unchanged.

Wires the bound through the worker's rendered config with the `max_new_per_run`
key the other automations already use, documents it in SKILL.md and README.md,
bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs.

Part of #656. Stacked on #654.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit cb1b625)

* chore(review): add read-only bounded-intake canary for #656

Loads the shipped catalog bundle and drives the real scheduled scan and
shared-intake drain against the live GitHub API. It starts no agent and posts
no comment: the dispatcher and the gate comment upsert record what would have
happened, so the canary is read-only. Development-only, kept in .pr/.

Co-authored-by: openhands <openhands@all-hands.dev>
(cherry picked from commit 9396c24)

* chore(review): drop the one-off bounded-intake canary harness

The 197-line read-only canary was evidence for #656, not part of the shipped
automation, so it does not belong in this focused per-scan bound PR. The live
result stays recorded in the PR description; the harness itself is preserved on
factory/reviewer-continuous-canary.

Co-authored-by: openhands <openhands@all-hands.dev>

---------

Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
@neubig
neubig force-pushed the fix/653-scheduled-review-retry branch from 0f3a2df to a8936f4 Compare September 25, 2026 20:55

@neubig neubig left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed after restacking #654 and the already-approved #659 layer onto current main. The scheduled retry, exact-head gating, bounded intake, generated catalogs, and focused regression suite are coherent; 219 focused tests pass locally and the generated indexes are clean.

@neubig
neubig merged commit c7e3a7d into main Sep 25, 2026
14 checks passed
@neubig
neubig deleted the fix/653-scheduled-review-retry branch September 25, 2026 20:56
@openhands-release-bot openhands-release-bot Bot added the released: v0.25.0 Shipped in v0.25.0 label Sep 27, 2026
@openhands-release-bot

Copy link
Copy Markdown
Contributor

🚀 Released in v0.25.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

released: v0.25.0 Shipped in v0.25.0 type: fix A bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants