Summary
A run_if gate that fails open is indistinguishable, from the job's point of view, from "nothing
changed". review-monitor's prci_delta gate has failed open on every pass it has ever logged
since 2026-09-11 (19/19), so merge/close detection for the 851 tasks in on_review has made zero
progress for 12 days, and nothing anywhere reported it. I found it by reading the artifact at the
top of a tick, not from any alert.
Evidence
/worktrees/.state/revmon-gate.log holds 19 lines. All 19 are the same verdict:
2026-09-11T04:05:01Z ... full_walk=yes tasks=0 elapsed=301.1s reason="gate_error: TimeoutError: "
...
2026-09-23T04:05:01Z ... full_walk=yes tasks=0 elapsed=301.1s reason="gate_error: TimeoutError: "
Elapsed is pinned at 300.0-306.6s across all 19, i.e. the timer fired as designed rather than the
work genuinely varying.
Why it cannot self-heal
overall_timeout_s defaults to 300.0 (prci_delta.py:340) and is armed around the whole run
(:1028, asyncio.wait_for(self._evaluate(...), timeout=self.overall_timeout_s or None)).
review-monitor's run_if block sets no overall_timeout_s, so the 300s default applies to the
largest queue in the fleet (851 rows; it was 621 when this was first analysed on 2026-09-12).
- The fail-open handler (
:1030-1033) builds _Outcome(job_id=..., satisfied=True, full_walk=True, reason=...) with no state= argument. _Outcome.state defaults to None,
documented None -> do not touch the state file.
- So
revmon-state.json is never written. It has been absent for 12 days. With no state, the next
pass must bootstrap all 851 deliverables, which is the most expensive pass possible, which
exceeds the cap, which fails open, which leaves the state absent. The loop is closed.
satisfied=True on that path is what makes it silent: nerve sees a satisfied gate and runs the job
normally. The job then reads full_walk: true and correctly refuses to transition anything, because
acting on a plan the gate could not compute is how the predicate loop this gate was built to end
gets recreated. Every one of the 19 ticks was a correct no-op, and that is precisely the loss.
A prior investigation on 2026-09-12 measured the same work at 86.7s and 111.5s with the cap lifted,
and fitted production per-call latency at roughly 27s fixed + 2.0s per gh call, which puts a
200-call pass near 430s. The cap sits between the idle and the loaded cost, which is why it is
100% rather than occasional. It also named 9 concrete transitions that were already frozen at that
point (8 merged PRs never advanced, 1 closed-unmerged never returned for a re-decision).
What I cannot do
config/cron/gates/prci_delta.py and config/cron/jobs.yaml are not in this repository (404) and
the workspace on the box is not a git repository, so propose_config_change refuses with "nothing
to open a pull request against". There is no reviewable surface for a one-value config change, and
my own rules forbid editing a gate or toolkit from inside a monitor tick. Six prior internal
attempts to route this through the normal implementation pipeline were all stopped for the same
reason, so I am asking instead of trying a seventh.
Which of these should I treat as the fix?
- A you set
overall_timeout_s: 1200 in review-monitor's run_if block in
config/cron/jobs.yaml. It is already a supported per-job float (_FLOAT_KEYS, :354), so this
is a value, not a code change.
- B you set
PRCI_OVERALL_TIMEOUT_S=1200 in the fleet unit environment
(ENV_PREFIX = "PRCI_", :318, read at :938). Same effect, but it applies to both
prci_delta jobs.
- C a change here so this class is not silent: either surface a gate that has failed open on N
consecutive passes, or let a bootstrap persist partial acks so the first pass after a state reset
is not all-or-nothing.
A or B unblocks the 851-task queue immediately; C is the part that stops the next gate from doing
this unnoticed. One word is enough and I will act on it.
Two caveats on A and B, so the result is not a surprise: the first pass that actually completes will
plan a large one-time wave of re-validation bounces (the module documents 459 of 586 on a previous
population), which is expected and not a malfunction. And /worktrees is instance store, so an EC2
stop/start wipes the state file and re-enters this same loop unless the cap is raised.
Summary
A
run_ifgate that fails open is indistinguishable, from the job's point of view, from "nothingchanged".
review-monitor'sprci_deltagate has failed open on every pass it has ever loggedsince 2026-09-11 (19/19), so merge/close detection for the 851 tasks in
on_reviewhas made zeroprogress for 12 days, and nothing anywhere reported it. I found it by reading the artifact at the
top of a tick, not from any alert.
Evidence
/worktrees/.state/revmon-gate.logholds 19 lines. All 19 are the same verdict:Elapsed is pinned at 300.0-306.6s across all 19, i.e. the timer fired as designed rather than the
work genuinely varying.
Why it cannot self-heal
overall_timeout_sdefaults to300.0(prci_delta.py:340) and is armed around the whole run(
:1028,asyncio.wait_for(self._evaluate(...), timeout=self.overall_timeout_s or None)).review-monitor'srun_ifblock sets nooverall_timeout_s, so the 300s default applies to thelargest queue in the fleet (851 rows; it was 621 when this was first analysed on 2026-09-12).
:1030-1033) builds_Outcome(job_id=..., satisfied=True, full_walk=True, reason=...)with nostate=argument._Outcome.statedefaults toNone,documented
None -> do not touch the state file.revmon-state.jsonis never written. It has been absent for 12 days. With no state, the nextpass must bootstrap all 851 deliverables, which is the most expensive pass possible, which
exceeds the cap, which fails open, which leaves the state absent. The loop is closed.
satisfied=Trueon that path is what makes it silent: nerve sees a satisfied gate and runs the jobnormally. The job then reads
full_walk: trueand correctly refuses to transition anything, becauseacting on a plan the gate could not compute is how the predicate loop this gate was built to end
gets recreated. Every one of the 19 ticks was a correct no-op, and that is precisely the loss.
A prior investigation on 2026-09-12 measured the same work at 86.7s and 111.5s with the cap lifted,
and fitted production per-call latency at roughly 27s fixed + 2.0s per
ghcall, which puts a200-call pass near 430s. The cap sits between the idle and the loaded cost, which is why it is
100% rather than occasional. It also named 9 concrete transitions that were already frozen at that
point (8 merged PRs never advanced, 1 closed-unmerged never returned for a re-decision).
What I cannot do
config/cron/gates/prci_delta.pyandconfig/cron/jobs.yamlare not in this repository (404) andthe workspace on the box is not a git repository, so
propose_config_changerefuses with "nothingto open a pull request against". There is no reviewable surface for a one-value config change, and
my own rules forbid editing a gate or toolkit from inside a monitor tick. Six prior internal
attempts to route this through the normal implementation pipeline were all stopped for the same
reason, so I am asking instead of trying a seventh.
Question (@alex-clickhouse)
Which of these should I treat as the fix?
overall_timeout_s: 1200inreview-monitor'srun_ifblock inconfig/cron/jobs.yaml. It is already a supported per-job float (_FLOAT_KEYS,:354), so thisis a value, not a code change.
PRCI_OVERALL_TIMEOUT_S=1200in the fleet unit environment(
ENV_PREFIX = "PRCI_",:318, read at:938). Same effect, but it applies to bothprci_deltajobs.consecutive passes, or let a bootstrap persist partial acks so the first pass after a state reset
is not all-or-nothing.
A or B unblocks the 851-task queue immediately; C is the part that stops the next gate from doing
this unnoticed. One word is enough and I will act on it.
Two caveats on A and B, so the result is not a surprise: the first pass that actually completes will
plan a large one-time wave of re-validation bounces (the module documents 459 of 586 on a previous
population), which is expected and not a malfunction. And
/worktreesis instance store, so an EC2stop/start wipes the state file and re-enters this same loop unless the cap is raised.