Skip to content

Cron gates that fail open are silent: review-monitor's prci_delta gate has timed out on 19/19 passes since 2026-09-11, freezing 851 tasks #454

Description

@groeneai

Summary

A run_if gate that fails open is indistinguishable, from the job's point of view, from "nothing
changed". review-monitor's prci_delta gate has failed open on every pass it has ever logged
since 2026-09-11 (19/19)
, so merge/close detection for the 851 tasks in on_review has made zero
progress for 12 days, and nothing anywhere reported it. I found it by reading the artifact at the
top of a tick, not from any alert.

Evidence

/worktrees/.state/revmon-gate.log holds 19 lines. All 19 are the same verdict:

2026-09-11T04:05:01Z ... full_walk=yes tasks=0 elapsed=301.1s reason="gate_error: TimeoutError: "
...
2026-09-23T04:05:01Z ... full_walk=yes tasks=0 elapsed=301.1s reason="gate_error: TimeoutError: "

Elapsed is pinned at 300.0-306.6s across all 19, i.e. the timer fired as designed rather than the
work genuinely varying.

Why it cannot self-heal

  1. overall_timeout_s defaults to 300.0 (prci_delta.py:340) and is armed around the whole run
    (:1028, asyncio.wait_for(self._evaluate(...), timeout=self.overall_timeout_s or None)).
  2. review-monitor's run_if block sets no overall_timeout_s, so the 300s default applies to the
    largest queue in the fleet (851 rows; it was 621 when this was first analysed on 2026-09-12).
  3. The fail-open handler (:1030-1033) builds _Outcome(job_id=..., satisfied=True, full_walk=True, reason=...) with no state= argument. _Outcome.state defaults to None,
    documented None -> do not touch the state file.
  4. So revmon-state.json is never written. It has been absent for 12 days. With no state, the next
    pass must bootstrap all 851 deliverables, which is the most expensive pass possible, which
    exceeds the cap, which fails open, which leaves the state absent. The loop is closed.

satisfied=True on that path is what makes it silent: nerve sees a satisfied gate and runs the job
normally. The job then reads full_walk: true and correctly refuses to transition anything, because
acting on a plan the gate could not compute is how the predicate loop this gate was built to end
gets recreated. Every one of the 19 ticks was a correct no-op, and that is precisely the loss.

A prior investigation on 2026-09-12 measured the same work at 86.7s and 111.5s with the cap lifted,
and fitted production per-call latency at roughly 27s fixed + 2.0s per gh call, which puts a
200-call pass near 430s. The cap sits between the idle and the loaded cost, which is why it is
100% rather than occasional. It also named 9 concrete transitions that were already frozen at that
point (8 merged PRs never advanced, 1 closed-unmerged never returned for a re-decision).

What I cannot do

config/cron/gates/prci_delta.py and config/cron/jobs.yaml are not in this repository (404) and
the workspace on the box is not a git repository, so propose_config_change refuses with "nothing
to open a pull request against". There is no reviewable surface for a one-value config change, and
my own rules forbid editing a gate or toolkit from inside a monitor tick. Six prior internal
attempts to route this through the normal implementation pipeline were all stopped for the same
reason, so I am asking instead of trying a seventh.

Question (@alex-clickhouse)

Which of these should I treat as the fix?

  • A you set overall_timeout_s: 1200 in review-monitor's run_if block in
    config/cron/jobs.yaml. It is already a supported per-job float (_FLOAT_KEYS, :354), so this
    is a value, not a code change.
  • B you set PRCI_OVERALL_TIMEOUT_S=1200 in the fleet unit environment
    (ENV_PREFIX = "PRCI_", :318, read at :938). Same effect, but it applies to both
    prci_delta jobs.
  • C a change here so this class is not silent: either surface a gate that has failed open on N
    consecutive passes, or let a bootstrap persist partial acks so the first pass after a state reset
    is not all-or-nothing.

A or B unblocks the 851-task queue immediately; C is the part that stops the next gate from doing
this unnoticed. One word is enough and I will act on it.

Two caveats on A and B, so the result is not a surprise: the first pass that actually completes will
plan a large one-time wave of re-validation bounces (the module documents 459 of 586 on a previous
population), which is expected and not a malfunction. And /worktrees is instance store, so an EC2
stop/start wipes the state file and re-enters this same loop unless the cap is raised.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions