Skip to content

Restore finished remote loops a host reboot killed - #501

Merged
scgopi merged 3 commits into
mainfrom
fix/codespace-dial-forever
Sep 29, 2026
Merged

scgopi merged 3 commits into
mainfrom
fix/codespace-dial-forever

Conversation

@scgopi

@scgopi scgopi commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

After a codespace or remote-host restart, a finished unattended loop (goal or time, state succeeded/failed/stalled/stopped) whose pane was still open dialed reboot wait-daemon indefinitely. It dialed every ~19s before 0.1.76, and since 0.1.76 it pauses after the #480 schedule, but it never came back on its own. Sibling loops on the same codespace recovered fine.

Root cause

  • The pane's reconnect decides who restores a missing session from the loop type alone (GhosttyTerminalView+Remote.swift:132-181). For an unattended loop it never launches: it logs reboot wait-daemon, exits 255 and waits for graphcoded.
  • The daemon's only restorer after a remote reboot is the 60s sweep GraphStore.ensureUnattendedSessionsAlive, which ensured runsUnattended && !isResolved only. isResolved is succeeded/failed/stalled/stopped (LoopNode.swift:574).
  • So a finished loop whose session was still alive at the restart (for example a session answering a follow-up, inside the resolved-session grace, or a stopped loop that was not killed) had a pane waiting on a daemon that would never ensure it.
  • The siblings that "recovered" in the reported dials.log were resumed by a human opening them: openNode sends .resumeSession for a resolved, non-stopped loop (AppFeature.swift:746-752). Every ensure resume in the log follows that session's connect wait-daemon within 0-13s. The sweep produced no ensure at all in two days.
  • 0.1.76 only caps the dialing (Codespaces ssh retries improvement #480 schedule, SSHReconnectLoop.codespaceScript). It does not change who restores the session. No other uncapped codespace dial path remains in the pane: codespaceScript retries on exit 1 and 255 under the schedule, and every other exit passes through.

Fix: one owner per state

Loop Restorer after a remote reboot Changed here
Attended (turn, sketch, composite) pane restoreScript no
Unattended, unresolved daemon ensure (sweep) now also records the boot marker
Unattended, resolved daemon sweep, resume-only new
  • The sweep hands finished loops to ZmxSessionLauncher.restoreRebootedRemote. That dials only when a pane of the host has redialed since the host's last answered probe. Each remote pane touches a per-host stamp (redialStamp) before every redial, and on a healthy host no pane redials, so the sweep spends nothing. Then one probe per host names the sessions that are missing and whose boot marker is from an earlier boot. Only those get a create dial, behind the same boot gate. A failed probe leaves the redial pending, so the host is probed once it is back even if every pane has paused.
  • The restore launches GraphStore.rebootRestoreCopy: the banked conversation via each backend's resume argv, with no prompt. If nothing is banked, it is a fresh session that opens on a short note, not the loop's task. It does not re-arm goal pollers or heartbeats, and it does not write node state.
  • The daemon's ensure now writes the boot marker when it creates a session and on each alive tick. Before, only a pane attach wrote it, so a session no pane had joined had no marker, and a daemon-restored one kept a stale one. remoteKillInvocation removes the marker, so a session the daemon ended on purpose (a finished loop freed, a stop, a delete) is never restored, and its pane reads "ended".
  • No double launch:
    • The pane never launches an unattended loop.
    • The daemon never ensures an attended one.
    • Daemon dials stay serialised per node by RemoteEnsureGate.
    • The remote script remains an atomic check-or-run.
  • Backends: all five (Claude Code, Copilot CLI, Codex, OpenCode, pi) support resume. The prompt-less fallback applies to any of them only when no ID was banked, for example a session that never reached its first SessionStart.

Test plan

RED: xcodebuild -scheme graphcode build-for-testing on aebe3f5 (tests only) -> exit 65, RemoteSessionResumeTests does not compile: no onRestoreRebootedSessions, rebootRestoreCopy, rebootProbeScript, parseRebootProbe, onlyAfterReboot; the idle-cost tests (aHealthyHostIsNeverProbed, aPaneStampsItsHostBeforeEveryRedial) do not compile on 8a3b5ad: no RebootProbeGate, no redialStamp
GREEN: xcodebuild -scheme graphcode test -only-testing RemoteSessionResumeTests, RemoteRebootRestoreTests, CodespaceDialScheduleTests, DialLogTests -> 48 tests in 4 suites passed, exit 0
REGRESSION: xcodebuild -scheme graphcode test (full) -> 1997 tests in 213 suites passed, exit 0; make check -> exit 0; graphcoded and graphcode-cli schemes build -> exit 0

Loopback rig: user-mode sshd with a fake HOME and a private ZMX_DIR, an isolated graphcoded, and one finished and one running goal loop. The reboot is simulated by killing both sessions and writing a stale boot marker. The pane side is the reconnect decision from GhosttyTerminalView+Remote.swift run over ssh.

Build Finished loop Running loop
0.1.76 daemon 14× reboot wait-daemon over 3.5 min and three sweeps, never ensured ensure fresh → reconnect attach-live
this branch 4× reboot wait-daemon, then ensure resume at the first sweep → reconnect attach-live; argv --resume transcript-of-finished, no goal text, state still succeeded ensure fresh → reconnect attach-live

On this branch, later sweeps logged nothing further: one agent launch per loop, one zmx session per loop.

Idle cost (codespace, 3 finished loops, no pane redialing)

One healthy gh codespace ssh spawn is 1 call from the codespaces bucket (GH_DEBUG=api, gh 2.87.0). The dial counts below were measured on the loopback rig as ssh sessions. The per-hour codespace figure is a model: dials multiplied by 1 call.

Build Measured idle dials Codespace API calls per hour (model)
main / 0.1.76 0 (finished loops skipped) 0
#501 first push (8a3b5ad) 3 probes in 4 min ~55-60
this revision 0 in 4 min, and 0 in 5 min on a second graph 0

Reboot on the revision, 3 finished loops with their sessions killed and markers from an earlier boot, store loaded:

Phase Probes Restores Other
No pane redial, 135s 0 0 0 sessions
After one pane redial stamp, 75s 1 3 3 agent launches, none carrying the loop task; all still succeeded
Next two sweeps 0 0 0 ssh sessions

With nothing banked, the fresh launch opens on the daemon's usual wake-memory pointer and the finished-loop note. This is the same shape the existing open-a-finished-loop path produces (resumeResolvedSession), never the loop's task.

During an outage, probes are bounded to one per sweep while panes redial, and panes stop on #480's schedule (4 min, then a pause). After that there is one restore dial per rebooted finished loop.

Known limits

  • Loops inside a composite's sub-graph are not swept on a remote host, before or after this change.
  • A restored finished session is not re-armed for the resolved-session end until the daemon next loads the graph.

Checklist

  • I have read the Contributing Guidelines
  • I have signed off my commits (git commit -s) per the DCO
  • Full macOS test suite and make check pass locally
  • I added the tests before the implementation and observed the RED failure

Signed-off-by: scgopi <scgopireddy@gmail.com>
A finished goal or time loop whose pane was open when its codespace
restarted dialed 'waiting for graphcoded' forever: the pane leaves every
unattended loop to the daemon, and the remote sweep skipped every resolved
node. The sweep now probes each host once for finished loops whose session
is missing and was last seen in an earlier boot, and brings those back as
their banked conversation, with no task, poller, heartbeat or state change.
The daemon's ensure records the boot marker too, and a daemon kill clears it,
so a session ended on purpose is never restored.

Signed-off-by: scgopi <scgopireddy@gmail.com>
The reboot probe ran on every liveness sweep once a project had a finished
loop, so an idle codespace went from no dials to one gh run a minute. A
remote pane now touches a per-host stamp before each redial, and the sweep
probes a host only when a stamp is newer than its last answered probe: a
healthy host costs nothing, and an outage is bounded by the pane's own

Signed-off-by: scgopi <scgopireddy@gmail.com>
#480 schedule.
@scgopi
scgopi merged commit b9272e9 into main Sep 29, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant