fix(review): gate agent launch on exact-head checks - #651
Conversation
The reviewer dispatched an agent conversation for every eligible trigger without reading the current head's check state, so a head with an already failing required check still spent an LLM slot (OpenHands/OpenHands#17544 created conversation f46343fd despite a failing Validate PR description). Add a deterministic head-eligibility gate in front of dispatcher.deliver: - GitHubRepository.check_runs(sha) reads the check runs the reviewer's own token can already see, with no ruleset or branch-protection access. - A completed run on the exact head blocks unless its conclusion is success, neutral, or skipped, so an unknown conclusion fails closed. - A queued or in_progress run exits as waiting on checks without holding a slot to poll, and never approves. - Runs attributed to another head are ignored. - The trigger is not consumed: the next scan or explicit review request starts the review once the head is green. - One explanation is left per head, upserted through a marker carrying the head SHA and gate category, and an equivalent repository workflow remediation comment is not duplicated. Bump the bundle version to 1.5.0 and regenerate the catalog artifacts. Co-authored-by: openhands <openhands@all-hands.dev>
|
@all-hands-bot please rename this PR to reflect the implemented policy: the gate rejects non-green exact-head checks, not only checks marked required. A title such as "fix(review): gate agent launch on exact-head checks" would match the issue decisions and the code. |
|
Renamed. The PR title is now: fix(review): gate agent launch on exact-head checks That matches the implemented policy: the gate classifies the exact head SHA's check runs, ignoring runs attributed to any other head, and blocks on any completed run whose conclusion is not Generated by OpenHands AI on behalf of the user. |
|
@all-hands-bot please tighten the deterministic gate before human review:
Please add narrow regression tests for the two ownership/deduplication cases and keep the patch focused. The current tests and CI are green, but these are live-path correctness gaps. |
|
@all-hands-bot I found a live edge case before rollout. GitHub's check-runs API lists both a failed and a later successful Please investigate the API data and add a regression test for repeated runs of the same check on one SHA. I have asked the maintainer whether the gate should use only the latest run per check name; hold that behavior change until the policy decision arrives. This is a live-path blocker for deploying #651 as written. |
Only treat a marked comment as managed when the configured reviewer account authored it, so a marker placed by a PR author can neither suppress the gate explanation nor be PATCHed. Narrow the workflow dedupe to a disclosure comment that names the same current-head checks. Rename the blocked heading to current-head checks. Add narrow regression tests for the ownership and dedupe cases.
|
Done on all three, plus narrow regression tests. Summary of the patch to 1. Ownership of managed comments. Added managed = [
comment
for comment in comments
if CHECK_GATE_MARKER in (comment.get("body") or "")
and self._owns_comment(comment)
]2. Narrower workflow dedupe. if any(
WORKFLOW_DISCLOSURE.lower() in (comment.get("body") or "").lower()
and self._explains_checks(comment.get("body") or "", names)
for comment in comments
if CHECK_GATE_MARKER not in (comment.get("body") or "")
):
return None3. Heading. The blocked heading is now Tests (
I did not push commits, since I only push when explicitly asked; say the word and I will push this to Separately, I see the follow-up about GitHub listing both a failed and a later successful Generated by OpenHands AI on behalf of the user. |
|
@all-hands-bot please push the three tested fixes you described (managed-comment ownership, check-specific workflow dedupe, and accurate heading) to this PR now. Leave the repeated-run behavior unchanged while the latest-run policy decision is pending. After pushing, report the new head SHA and current CI results. Do not merge. |
|
Pushed. The three tested fixes are now on this PR branch, and I did not merge anything. New head SHA: What landed:
No change to repeated-run behavior: Current CI results on
Locally, Generated by OpenHands AI on behalf of the user. |
|
Deployment blocker found during live audit: the OSS reviewer automation has only |
|
Policy decision for the repeated-run blocker: classify only the latest run of each logical check on the exact head. Use check name plus reporting app identity as the grouping key, and deterministic run start time/ID ordering; a later queued/in-progress rerun is waiting, a later success supersedes an earlier failure on the same SHA. This matches the live |
|
@all-hands-bot please implement the latest-run policy recorded above on this PR now, using the live #647/#649 repeated |
|
One edge case in head |
GitHub lists every run for a commit, so a check re-run after a fix kept contributing its superseded failure forever. Group check runs by name and reporting app identity, pick the latest by start time and then run ID, and classify only that run. A later queued or in-progress re-run supersedes an earlier success, so the head waits instead of being reviewed. Adds regression tests for the live #647/#649 repeated Validate PR description runs, the start-time tie broken by run ID, a newer pending re-run, and same-name checks from different apps. Co-authored-by: openhands <openhands@all-hands.dev>
5261e28 to
336aaf2
Compare
|
Implemented the latest-run policy you recorded and pushed it. I did not merge. New head SHA: What changed
Regression testsAdded to
Verified against the live dataI ran the shipped gate methods against the actual check-run payloads for both heads, rather than trusting the fixtures alone. Both now classify green where they previously blocked forever, and the kept run is the later success: In each case the superseded failure ( Local and CI results on
|
| Check | Workflow | Result |
|---|---|---|
| test | Tests | ✅ success |
| validate-claude-code | Check Extensions | ✅ success |
| sync-extensions | Check Extensions | ✅ success |
| sync-sdk-skill | Check Extensions | ✅ success |
| check | Deprecation deadlines | ✅ success |
| Validate PR description | PR Description Check | ✅ success |
| check-pr-artifacts | PR Artifacts | ✅ success |
| pr-title / Lint PR title (conventional) | pr | ✅ success |
| pr-title / Apply conventional-type label | pr | ✅ success |
| cleanup-on-approval | PR Artifacts | ⏭️ skipped |
| cleanup-after-merge | PR Artifacts | ⏭️ skipped |
Merge state is BLOCKED pending review and the branch is MERGEABLE; nothing was merged. The pending-to-green retry gap you described stays out of this PR for #653 to stack on once this base is green.
Generated by OpenHands AI on behalf of the user.
|
@all-hands-bot please address the queued-run edge case in my latest comment before this check gate is reviewed or deployed. Add a regression test for a newer queued run with no started_at and push the focused fix; do not merge. |
A queued or requested check run can exist before GitHub sets started_at, which the API reports as null. Ordering used the empty string for a missing start time, so an earlier completed success with a timestamp outranked the newer queued rerun and the reviewer could launch while the rerun was still pending. Fall back to the run ID, which increases monotonically, whenever a start time is absent. Co-authored-by: openhands <openhands@all-hands.dev>
|
The new |
|
Fixed and pushed. I did not merge. New head SHA: The bugYou are right, and it is exactly as you described. return (run.get("started_at") or "", int(run.get("id") or 0))A queued/requested run that has no The earlier success won, the queue was classified green, and the reviewer launched while the rerun was still pending. The fixrun_id = int(run.get("id") or 0)
started_at = run.get("started_at")
return (started_at or f"~{run_id:020d}", run_id)When a start time is absent, the run ID (monotonic per repository) fills in as the primary key, so a newer queued run always outranks an earlier completed run on the same logical check. Timestamped ordering is unchanged, and a timestamp tie still falls through to the run ID as before. I re-checked the other orderings, and they all still hold: later timestamp beats an earlier one regardless of ID; two runs with no start time order by ID; two runs sharing a timestamp order by ID. Regression test
Verification on
|
| Check | Workflow | Result |
|---|---|---|
| test | Tests | ✅ pass |
| validate-claude-code | Check Extensions | ✅ pass |
| sync-extensions | Check Extensions | ✅ pass |
| sync-sdk-skill | Check Extensions | ✅ pass |
| check | Deprecation deadlines | ✅ pass |
| Validate PR description | PR Description Check | ✅ pass |
| check-pr-artifacts | PR Artifacts | ✅ pass |
| pr-title / Lint PR title (conventional) | pr | ✅ pass |
| pr-title / Apply conventional-type label | pr | ✅ pass |
| cleanup-on-approval | PR Artifacts | ⏭️ skipping |
| cleanup-after-merge | PR Artifacts | ⏭️ skipping |
Merge state is BLOCKED pending review and the branch is MERGEABLE; nothing was merged, per your instruction. The change is scoped to the ordering helper, its docs, and the one regression test — the pending-to-green retry gap stays in #653 to stack on this base.
Generated by OpenHands AI on behalf of the user.
|
@all-hands-bot please fix the reverse-order case in my newest comment and add the regression. The current |
The `~{run_id}` fallback sorted a run without a `started_at` past every
real timestamp, so an older queued/cancelled run could outrank a newer
successful run on the same check and keep a head blocked or waiting. The
check-run ID is the monotonic creation sequence, so make it the primary
order key and use the start time only as a tie-break. Adds the reverse
regression alongside the existing null-start case.
|
Fixed and pushed. I did not merge. New head SHA: The bugYou are right. run_id = int(run.get("id") or 0)
started_at = run.get("started_at")
return (started_at or f"~{run_id:020d}", run_id)
The fixThe run ID is the reliable creation sequence, so it is now the primary key and the start time only breaks a tie: run_id = int(run.get("id") or 0)
return (run_id, run.get("started_at") or "")I verified both directions against the old and new keys directly:
The existing Regression
Verification on
|
| Check | Workflow | Result |
|---|---|---|
| test | Tests | ✅ pass |
| validate-claude-code | Check Extensions | ✅ pass |
| sync-extensions | Check Extensions | ✅ pass |
| sync-sdk-skill | Check Extensions | ✅ pass |
| check | Deprecation deadlines | ✅ pass |
| Validate PR description | PR Description Check | ✅ pass |
| check-pr-artifacts | PR Artifacts | ✅ pass |
| pr-title / Lint PR title (conventional) | pr | ✅ pass |
| pr-title / Apply conventional-type label | pr | ✅ pass |
| cleanup-on-approval | PR Artifacts | ⏭️ skipping |
| cleanup-after-merge | PR Artifacts | ⏭️ skipping |
Nothing was merged, and the pending-to-green retry remains in #653 to stack on this base.
Generated by OpenHands AI on behalf of the user.
The OSS reviewer's only trigger was the GitHub review_requested event. #651 stopped dispatching an agent while the exact head's checks were pending or failing, and its waiting comment promised a scheduled retry that deployment did not have, so a request arriving during CI could be left unreviewed. Reuse the existing worker and one automation record. In scheduled mode the scan now also considers every open, non-draft PR that still holds an outstanding all-hands-bot review request, keyed by that request's own review_requested event so repeated scans reuse one conversation and one review. The explicit request event path is unchanged for event-only deployments. The waiting and blocked comments now name the retry the deployment actually has: a scheduled scan where a cron trigger exists, and another review request where only the event trigger does. Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and adds scheduled-scan, duplicate-scan, draft, and gate-message tests. Part of #653. Stacked on #651. Co-authored-by: openhands <openhands@all-hands.dev>
|
@all-hands-bot Please update this PR to cover a current-head GitHub Actions workflow that fails before creating any check runs. This is within #644's existing goal of not launching a reviewer for red CI. Concrete reproduction: #426 head Please make the pre-dispatch gate consider current-head workflow-run conclusions as well as check runs, with the same blocked/waiting behavior and a focused regression test. Ignore obsolete-head runs, and avoid duplicating the existing gate/comment machinery. This should be part of #651 rather than a separate PR. |
The head-eligibility gate read only check runs, so a workflow that failed before creating any check run left the commit's check-run rollup green and the reviewer launched for red CI. On #426 head 41ffb9d, `Tests`, `Check Extensions`, and `Deprecation deadlines` failed with zero jobs, so only the green `pr-title` checks were visible; the reviewer spent an agent slot before reporting the failures, and `gh pr checks` omitted the zero-job runs as well. Read the current head's Actions workflow runs as well as its check runs: - GitHubRepository.workflow_runs(sha) reads `/actions/runs?head_sha=<sha>`, which answers with an object and so paginates manually like check_runs. - A workflow run whose check suite already reported check runs is left to those runs, so a workflow is never counted twice. Only a run whose suite reported no check runs - a workflow-level failure, or a `pull_request` run whose jobs never started - is added to the gate. - Workflow runs get the same blocked/waiting behavior and the same current-head filter as check runs, and the latest run of each workflow (name + workflow ID, by run ID then start time) supersedes its superseded attempts. - The two grouping paths share one `_latest_by_group` helper and the one existing classifier and comment machinery. Co-authored-by: openhands <openhands@all-hands.dev>
|
Fixed and pushed. I did not merge. New head SHA: The gapYou are right, and your reproduction is exact. The gate read only The change
Regression test
Live verificationRunning the shipped The reproduced head is now blocked on the three zero-job workflows, and this PR's own head still reads green (its workflows all created check runs, so they are not double-counted).
Nothing was merged. Generated by OpenHands AI on behalf of the user. |
|
@all-hands-bot The workflow-run gate fix is pushed, but this PR now reports CONFLICTING with main. Please resolve only the merge conflict, preserve the exact-head check-run and workflow-run gate behavior, run focused tests plus CI, and push the branch. Do not merge. |
Resolve the CONFLICTING state reported on #651. The only conflicted files were the generated bundles (skills/index.js, automations/bundle-index.js), which both sides had regenerated; they are resolved by re-running `npm run build` over the auto-merged sources, so both the workflow-run gate from this branch and main's refreshed-state prompt/docs are preserved. Gate behavior is unchanged: exact-head check-run and workflow-run classification, latest-run-per-logical-check ordering, owned-comment and check-specific dedupe.
|
Conflict resolved and pushed. I did not merge. New head SHA: The conflictOnly two files conflicted, both generated:
Both sides had regenerated them — this branch from the workflow-run gate, I resolved the two generated files by taking the merged sources and re-running Gate behavior preservedThe merge is additive for this PR:
I also ran the shipped This PR's head still reads green, and the #426 head is still blocked on the three zero-job workflows, so the workflow-run gate survived the merge. Tests
CI on
|
| Check | Workflow | Result |
|---|---|---|
| test | Tests | ✅ pass |
| validate-claude-code | Check Extensions | ✅ pass |
| sync-extensions | Check Extensions | ✅ pass |
| sync-sdk-skill | Check Extensions | ✅ pass |
| check | Deprecation deadlines | ✅ pass |
| Validate PR description | PR Description Check | ✅ pass |
| check-pr-artifacts | PR Artifacts | ✅ pass |
| pr-title / Lint PR title (conventional) | pr | ✅ pass |
| pr-title / Apply conventional-type label | pr | ✅ pass |
| cleanup-on-approval | PR Artifacts | ⏭️ skipping |
| cleanup-after-merge | PR Artifacts | ⏭️ skipping |
Nothing was merged. The scope stayed at the conflict: I did not change gate policy or touch the #653 pending-to-green retry.
Generated by OpenHands AI on behalf of the user.
The OSS reviewer's only trigger was the GitHub review_requested event. #651 stopped dispatching an agent while the exact head's checks were pending or failing, and its waiting comment promised a scheduled retry that deployment did not have, so a request arriving during CI could be left unreviewed. Reuse the existing worker and one automation record. In scheduled mode the scan now also considers every open, non-draft PR that still holds an outstanding all-hands-bot review request, keyed by that request's own review_requested event so repeated scans reuse one conversation and one review. The explicit request event path is unchanged for event-only deployments. The waiting and blocked comments now name the retry the deployment actually has: a scheduled scan where a cron trigger exists, and another review request where only the event trigger does. Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and adds scheduled-scan, duplicate-scan, draft, and gate-message tests. Part of #653. Stacked on #651. Co-authored-by: openhands <openhands@all-hands.dev>
|
@all-hands-bot Please update the existing implementation in #651; do not open another PR. The live deployment exposed two over-broad parts of the gate:
Please update the focused tests and PR description accordingly. Add a regression fixture matching #17200: one successful required check plus an optional zero-job workflow |
The live deployment exposed two over-broad behaviors in the exact-head eligibility gate. OpenHands/OpenHands#17200 passes `test-and-build (ubuntu)`, its only GitHub-required check, while the optional `release ready` workflow has a zero-job `startup_failure`. Classifying every check and workflow run treated that optional failure as blocking, so scheduled discovery paused a mergeable head. An explicit `all-hands-bot` review request is the intake-policy exception: the caller asked for that head by name, so it must dispatch even when required CI is red or pending. - GitHubRepository.required_check_contexts(number) adds one GraphQL query that returns the head's status check rollup and keeps only the contexts isRequired(pullRequestNumber:) marks required. That field is merge-policy source of truth and is pull-request-scoped, so it is correct for a stacked PR whose symbolic base branch has no rules of its own. - Scheduled classification uses only those required contexts: a required check run by name, a required commit status by context. An optional failure cannot block. A required context with no current-head run is expected, so it waits. - When the required signal cannot be read or is empty the gate falls back to every current-head check run and workflow run, so a red head still blocks. This keeps the #426 zero-job shape blocked: its required checks never started, isRequired returns no node, and the empty set falls back. - _gate_head(pr, requested=True) returns green without reading checks and without a gate comment, and run() passes requested=event_mode. Exact-head filtering, latest-run supersession, fail-closed conclusions, the comment marker and its ownership/dedupe rules, and the conversation dispatch path are unchanged. Co-authored-by: openhands <openhands@all-hands.dev>
|
Updated and pushed. I did not merge. New head SHA: Both over-broad parts are fixed1. Scheduled discovery now classifies only GitHub-required checks. I used the signal you verified: 2. An explicit request bypasses the gate. Your
|
| required set | scheduled classification | |
|---|---|---|
#17200 (test-and-build (ubuntu) passes, optional release ready startup_failure) |
[{name: "test-and-build (ubuntu)", kind: "CheckRun"}] |
('green', []) → dispatches |
explicit review_requested request |
not read (bypass) | dispatches, no gate comment |
A regression fixture matching it, test_reviewer_ignores_an_optional_zero_job_workflow_on_a_green_required_head, asserts exactly one dispatch and no gate comment; test_reviewer_explicit_request_bypasses_the_ci_gate asserts the request dispatches with a failing required check and posts no comment. Both fail against the previous all-run classifier and event-gated code, and pass now.
The #426 case still blocks
isRequired is a field on each context that actually reported, so a required check that failed before creating any check run yields no node and the required set comes back empty. An empty or unreadable set falls back to the full current-head check run + workflow run rollup, so #426 still blocks rather than launching:
OpenHands/extensions 426 required -> []
OpenHands/extensions 426 scheduled -> ('blocked', ['Check Extensions', 'Deprecation deadlines', 'Tests'])
OpenHands/OpenHands 17200 scheduled -> ('green', [])
OpenHands/extensions 651 scheduled -> ('green', [])
That fallback is also the fail-closed path: a GraphQL error logs required_check_signal: unavailable and classifies the full rollup, so a required-check read failure never approves silently.
Preserved machinery
Exact-head filtering, latest-run-per-logical-check/workflow supersession (run ID primary, start time tie-break), fail-closed conclusions, the openhands-review-gate comment marker with its ownership filter and check-specific dedupe, and the conversation dispatch path are all unchanged. _classify_runs now holds the shared status/conclusion logic, with _classify_all_runs and _classify_required_runs as the two callers.
Verification
- Focused:
uv run --group test pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py— 60 passed. - Full suite:
uv run --group test pytest -q— 1021 passed, 23 skipped. scripts/sync_extensions.py --check— clean apart from the pre-existing non-blockingissue-duplicate-checkerwarning.npm run build— no drift.
CI on 9c62801 — all green (9 pass, 2 skipped, 0 failing, 0 pending)
| Check | Workflow | Result |
|---|---|---|
| test | Tests | ✅ pass |
| validate-claude-code | Check Extensions | ✅ pass |
| sync-extensions | Check Extensions | ✅ pass |
| sync-sdk-skill | Check Extensions | ✅ pass |
| check | Deprecation deadlines | ✅ pass |
| Validate PR description | PR Description Check | ✅ pass |
| check-pr-artifacts | PR Artifacts | ✅ pass |
| pr-title / Lint PR title (conventional) | pr | ✅ pass |
| pr-title / Apply conventional-type label | pr | ✅ pass |
| cleanup-on-approval | PR Artifacts | ⏭️ skipping |
| cleanup-after-merge | PR Artifacts | ⏭️ skipping |
The PR description is updated to match: required-only scheduled classification, the explicit-request exemption, the fail-closed fallback, and the new tests. Nothing was merged.
Generated by OpenHands AI on behalf of the user.
The OSS reviewer's only trigger was the GitHub review_requested event. #651 stopped dispatching an agent while the exact head's checks were pending or failing, and its waiting comment promised a scheduled retry that deployment did not have, so a request arriving during CI could be left unreviewed. Reuse the existing worker and one automation record. In scheduled mode the scan now also considers every open, non-draft PR that still holds an outstanding all-hands-bot review request, keyed by that request's own review_requested event so repeated scans reuse one conversation and one review. The explicit request event path is unchanged for event-only deployments. The waiting and blocked comments now name the retry the deployment actually has: a scheduled scan where a cron trigger exists, and another review request where only the event trigger does. Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and adds scheduled-scan, duplicate-scan, draft, and gate-message tests. Part of #653. Stacked on #651. Co-authored-by: openhands <openhands@all-hands.dev>
|
Live Canvas evidence (2026-09-23):
No merge was performed. |
VascoSch92
left a comment
There was a problem hiding this comment.
Accepting as from the evidences it seems to work fine
The OSS reviewer's only trigger was the GitHub review_requested event. #651 stopped dispatching an agent while the exact head's checks were pending or failing, and its waiting comment promised a scheduled retry that deployment did not have, so a request arriving during CI could be left unreviewed. Reuse the existing worker and one automation record. In scheduled mode the scan now also considers every open, non-draft PR that still holds an outstanding all-hands-bot review request, keyed by that request's own review_requested event so repeated scans reuse one conversation and one review. The explicit request event path is unchanged for event-only deployments. The waiting and blocked comments now name the retry the deployment actually has: a scheduled scan where a cron trigger exists, and another review request where only the event trigger does. Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and adds scheduled-scan, duplicate-scan, draft, and gate-message tests. Part of #653. Stacked on #651. Co-authored-by: openhands <openhands@all-hands.dev>
* fix(review): resume requested reviews on a scheduled check scan The OSS reviewer's only trigger was the GitHub review_requested event. #651 stopped dispatching an agent while the exact head's checks were pending or failing, and its waiting comment promised a scheduled retry that deployment did not have, so a request arriving during CI could be left unreviewed. Reuse the existing worker and one automation record. In scheduled mode the scan now also considers every open, non-draft PR that still holds an outstanding all-hands-bot review request, keyed by that request's own review_requested event so repeated scans reuse one conversation and one review. The explicit request event path is unchanged for event-only deployments. The waiting and blocked comments now name the retry the deployment actually has: a scheduled scan where a cron trigger exists, and another review request where only the event trigger does. Bumps the catalog entry and bundle to 1.6.0, documents the retry contract, and adds scheduled-scan, duplicate-scan, draft, and gate-message tests. Part of #653. Stacked on #651. Co-authored-by: openhands <openhands@all-hands.dev> * fix(review): reword the managed gate comment across a trigger change A pending PR first gated by an event-only deployment kept its 'request all-hands-bot again' comment after the automation was switched to a cron scan, because the gate returned early on a matching state:sha marker without comparing the body. The comment therefore named a retry that deployment no longer had. The gate now compares the full managed body for the same marker and rewrites the comment in place when it changed, so an event-only comment becomes the scheduled one once the scan exists and an identical body stays a no-op. No second comment is created. The event-only wording also makes the retry explicit: the outstanding request must be removed and re-requested, since GitHub will not accept a second request for a reviewer who is already requested. Adds a focused event-then-cron regression plus a stale-body rewrite test, and documents the in-place reword in SKILL.md and README.md. Part of #653. Co-authored-by: openhands <openhands@all-hands.dev> * fix(review): bound scheduled PR review intake across repositories (#659) * fix(review): bound scheduled PR review intake across repositories A scheduled scan resumed every outstanding all-hands-bot request it found and started one conversation for each, so a first scan over the OSS backlog could start an agent for hundreds of eligible pull requests at once and exhaust the deployment. The scan now starts at most `max_new_per_run` new review conversations, defaulting to 2, counted across every configured repository rather than reset per repository. Delivery reuses the existing machinery: eligible candidates are collected with the request event and head SHA #654 already keys deliveries on, ordered by the oldest outstanding request and then by repository and pull-request number, and started through the same `dispatcher.deliver`. The bound counts conversations a scan starts, so a deduplicated or already-running delivery reuses its runtime and consumes no slot, and a candidate whose exact-head checks are pending or failing posts its gate disposition and consumes none either. Reaching the bound still evaluates the remaining candidates and their gate comments, and a dispatch that raises is reported without aborting the scan. The explicit reviewer-request event path and the trigger-label scan are unchanged. Wires the bound through the worker's rendered config with the `max_new_per_run` key the other automations already use, documents it in SKILL.md and README.md, bumps the catalog entry and bundle 1.6.0 -> 1.7.0, and regenerates the catalogs. Part of #656. Stacked on #654. Co-authored-by: openhands <openhands@all-hands.dev> (cherry picked from commit cb1b625) * chore(review): add read-only bounded-intake canary for #656 Loads the shipped catalog bundle and drives the real scheduled scan and shared-intake drain against the live GitHub API. It starts no agent and posts no comment: the dispatcher and the gate comment upsert record what would have happened, so the canary is read-only. Development-only, kept in .pr/. Co-authored-by: openhands <openhands@all-hands.dev> (cherry picked from commit 9396c24) * chore(review): drop the one-off bounded-intake canary harness The 197-line read-only canary was evidence for #656, not part of the shipped automation, so it does not belong in this focused per-scan bound PR. The live result stays recorded in the PR description; the harness itself is preserved on factory/reviewer-continuous-canary. Co-authored-by: openhands <openhands@all-hands.dev> --------- Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com> Co-authored-by: openhands <openhands@all-hands.dev> --------- Co-authored-by: openhands <openhands@all-hands.dev> Co-authored-by: all-hands-bot <all-hands-bot@users.noreply.github.com>
|
🚀 Released in v0.25.0. |
HUMAN:
On 2026-09-22 a review request for
all-hands-botonOpenHands/OpenHands#17544created a reviewer conversation even though the requiredValidate PR descriptioncheck was already failing, spending an LLM slot on a head that deterministically could not be reviewed. Being conservative is also what the maintainer asked for afterOpenHands/OpenHands#16223was auto-approved over a failing current-head gate. The fix adds a deterministic check-run eligibility gate in front of agent dispatch. The evidence below is agent-run tests plus live API reads of the reported heads; I intentionally left the human-test checkbox unchecked because no human exercised the change.AGENT:
Why
skills/github-pr-reviewer/scripts/worker.pydispatched an agent conversation for every eligible trigger without looking at the current head's check state. On 2026-09-22, requestingall-hands-botonOpenHands/OpenHands#17544created conversationf46343fd-9486-5943-a38f-fd8a279dd40feven though the head's requiredValidate PR descriptioncheck was alreadyCOMPLETED/FAILURE. The reviewer should reserve agent work for heads that pass cheap deterministic eligibility checks.Iterating on the live deployment exposed the two failure modes the gate must not have:
OpenHands/extensions#426at head41ffb9dhad three failedpull_requestworkflow runs (Tests,Check Extensions,Deprecation deadlines), each with zero jobs, while the commit check-run rollup was green, so the reviewer launched anyway.OpenHands/OpenHands#17200passestest-and-build (ubuntu), its only GitHub-required check, while the optionalrelease readyworkflow has a zero-jobstartup_failure; the all-run classifier treated that optional failure as blocking.The gate now classifies on GitHub's required-check signal in scheduled mode, and the explicit-request intake path is exempt.
Summary
skills/github/scripts/github_client.pyGitHubRepository.check_runs(sha)reads/repos/{owner}/{repo}/commits/{sha}/check-runs. That endpoint answers with an object rather than a list, so it paginates manually rather than throughgh_pages.GitHubRepository.workflow_runs(sha)reads/repos/{owner}/{repo}/actions/runs?head_sha={sha}the same way. A workflow that fails before any job reports leaves a failed check suite with no check runs under it, so only the workflow run reveals it.GitHubRepository.required_check_contexts(number)runs one GraphQL query returning the head's status check rollup and keeps only the contextsisRequired(pullRequestNumber:)marks required. That field is the merge policy's own source of truth and is pull-request-scoped, so it stays correct for a stacked PR (#17200's symbolic base branch carries no branch rules) where a ruleset or branch-protection read would not.skills/github-pr-reviewer/scripts/worker.py— the head-eligibility gate, evaluated immediately beforedispatcher.deliver(...):commit statusresolves throughstatuses(). An optional workflow that fails - including one that fails before creating any check run - cannot block.all-hands-botreview request bypasses the gate. The caller asked for that head by name, so_gate_head(pr, requested=True)returns green without reading checks and without leaving a gate comment. Draft, scope, exact-head, and delivery-deduplication safeguards still apply.#426: its required checks never started, soisRequiredreturns no node and the empty set falls back to the all-run rollup.success,neutral, orskippedclassifies as blocked; a run whose status is notcompleted, or a required context that has not reported a current-head run at all, classifies as waiting. An unrecognized conclusion fails closed rather than approving silently, and an unfulfilled required check is never approval.review-blocked/review-waitingdisposition andcontinues before any conversation is created, without holding a worker slot to poll. An existing completed review still short-circuits first, so the gate never second-guesses a finished verdict.all-hands-botrequest, starts the review once the head's required checks are non-blocking._gate_comment()leaves exactly one concise explanation on the PR. It upserts its own comment through a stable hidden marker carrying the head SHA and gate category (<!-- openhands-review-gate:{blocked|waiting}:{sha} -->), so a later run for a different head updates the marked comment instead of duplicating it. Only a marker this reviewer account authored counts as its own; any other marker is untrusted. If the repository workflow already posted a deterministic remediation comment naming the same reported checks, the gate adds nothing.skills/github-pr-reviewer/SKILL.md,README.md— document the required-only scheduled gate, the explicit-request exemption, the blocked/waiting outcomes, and the retry contract.No change to the review prompt, verdict parsing,
scripts/maintainer_handoff.py, profile/secret selection, or the Actions-basedplugins/pr-review.Issue Number
Fixes #644
How to Test
uv run --group test pytest -q tests/test_github_reviewer_delivery.py tests/test_github_automation_foundation.py— expect all pass.uv run --group test pytest -q— expect 1021 passed, 23 skipped.uv run python scripts/sync_extensions.py --check— expect clean (the pre-existing non-blockingissue-duplicate-checkercoverage warning is unrelated).npm run build— expect no drift inautomations/bundle-index.jsorskills/index.js.Regression coverage in
tests/test_github_reviewer_delivery.pydrives the real shippedworker.pyentrypoint through the catalog bundle helper:test_reviewer_ignores_an_optional_zero_job_workflow_on_a_green_required_head— the#17200fixture: one requiredsuccesscheck plus an optional zero-jobstartup_failureworkflow yields exactly one dispatch and no gate comment.test_reviewer_explicit_request_bypasses_the_ci_gate— a nativereview_requestedevent targetingall-hands-botdispatches even with a failing required check, and leaves no gate comment.test_reviewer_blocks_when_a_required_check_fails— a failed required check blocks and names only the required check, not the optional one.test_reviewer_waits_for_a_pending_required_check,test_reviewer_waits_for_a_required_context_with_no_current_head_run— pending and unfulfilled required checks wait.test_reviewer_falls_back_to_all_checks_when_required_reported_nothing,test_reviewer_falls_back_to_every_check_when_required_signal_is_unavailable— the#426empty-required and signal-failure paths still block.test_reviewer_ignores_an_optional_failure_but_reads_required_statuses— a required commit status is classified alongside required check runs.test_reviewer_blocks_a_workflow_that_failed_with_no_check_runs,test_reviewer_does_not_double_report_a_workflow_that_has_check_runs,test_reviewer_waits_for_a_workflow_run_that_has_not_finished,test_reviewer_ignores_workflow_runs_from_an_obsolete_head,test_reviewer_lets_a_green_workflow_rerun_supersede_a_failed_one— the workflow-run fallback and its exact-head, dedupe, and supersession rules.test_reviewer_blocking_current_head_check_creates_no_conversation,test_reviewer_green_current_head_creates_exactly_one_conversation,test_reviewer_pending_current_head_check_waits_without_a_conversation,test_reviewer_ignores_checks_from_an_obsolete_head,test_reviewer_fails_closed_on_an_unknown_conclusion— the core gate contract.test_reviewer_does_not_duplicate_the_gate_comment_for_the_same_head,test_reviewer_updates_its_gate_comment_when_the_head_moves,test_reviewer_defers_to_an_equivalent_workflow_remediation_comment,test_reviewer_ignores_a_marked_comment_a_pr_author_wrote,test_reviewer_does_not_defer_to_a_workflow_comment_about_another_check— the comment-marker and dedupe contract.test_reviewer_waiting_head_does_not_consume_a_later_green_request— a waiting stop does not consume the trigger.tests/test_github_automation_foundation.pycovers the transport: paginated object responses forcheck_runsandworkflow_runs, andrequired_check_contextskeeping only required nodes, rejecting an over-long rollup, and surfacing GraphQL errors.Live API reads
Required-check signal through the new helper with the reviewer's own token:
Scheduled classification of the reported heads through the shipped classifier:
#17200's optionalrelease readystartup_failureno longer blocks;#426still falls back to the all-run rollup and blocks.Video/Screenshots
Not applicable: this changes automation worker dispatch logic, not a GUI.
Notes
_gate_commentreads the PR's existing comments on every blocked or waiting run. That is one paginated read on a path that already performs several, and it does not create an agent conversation.required_check_contextsis a single GraphQL query per candidate, and an explicit request skips it entirely.Generated by OpenHands AI on behalf of the user.