Skip to content

ios(runner): a 1 s readiness preflight abandons a busy runner command and the next connect stalls behind it, failing waits on cold hosts #2475

Description

@thymikee

Current disposition: reproduce the cause before changing preflight policy

Audited main e4716e6f514273f060b4f1fb8a523d3c51a9c17d on 2026-10-02. The symptom is real; the specific queue-blocking explanation and its five-second/one-second regression proof remain unestablished. Keep this issue open for a bounded verification pass, rather than count #2804 or #2982 as its completion.

Facts that constrain the reproduction:

  • Ready sessions still cap their uptime preflight at 1,000 ms. Starting read-only sessions skip that preflight; the exemption does not cover every read-only command.
  • At the issue's cited baseline f4c8f3ddda, handleRequestBody already served .uptime inline before commandExecutionQueue.async; current main retains that separation through inlineResponse. A fixture that queues uptime behind a five-second native command would invent a production behavior. Exercise the real transport and identify what actually remains busy if the scenario reproduces.
  • fix(ios-runner): resend read-only commands across the RUNNER_BUSY drain window #2804 extends read-only RUNNER_BUSY draining retries. fix(ios-runner): a caller deadline no longer kills a starting runner #2982 preserves a detached, finitely budgeted runner start after a caller deadline. Neither proves this historical scenario fixed.
  • Latest failed main iOS run 36916418541 had a Settings step-4 wait succeed on its second attempt. Its fixture failure was a deep-link destination wait with APP_NOT_RUNNING, 15 completed retriable polls and zero readable captures. Those symptoms do not establish abandoned-preflight blocking. Latest head's 36999101157 passed both Settings replay and fixture smoke; one green run does not meet the original repeated-cold-run criterion.

Bounded completion contract

  1. Reach a control over the real runner transport with a native command deliberately held busy for a modeled five-second interval. Measure uptime response, the one-second preflight deadline, cancellation, and the next command separately. Release the block deterministically; avoid a five-second wall-clock unit sleep. Demonstrate both a healthy control and the claimed failure before designing a fix.
  2. Reproduce the wait shape after a roughly five-second snapshot-source attempt, with its actual remaining deadline forwarded to the runner fallback. Preserve the deadline; do not widen smoke timeouts or retries.
  3. If the actual path fails, fix the owning budget/transport/runner seam and make the regression fail on the pre-fix implementation. An artificial serial fake uptime cannot establish that seam. Confirm fixture smoke and Settings replay on multiple consecutive cold main runs before claiming the incident resolved.
  4. If uptime remains responsive and the historic mechanism cannot be reproduced, record the commands, baseline, timings and native routing controls, then close this historical cause hypothesis explicitly. Any different recurring flake needs its own captured cause and owning fix; a generic wait timeout is insufficient evidence for this issue.

Historical report follows; its interpretation is a hypothesis, not a confirmed native execution trace.

Symptom

iOS live lanes (Run fixture-backed iOS simulator E2E smoke, Run iOS Settings replay smoke test) fail a wait that follows open with wait timed out for …, on unrelated branches and on main. Seen 2026-09-10 on main f4c8f3ddda (16:44), on #2473's branch (16:15, passed on retry), on #2471 (17:00), and earlier on 2026-09-06/07 on at least four unrelated branches.

What the request logs show

From the artifacts of runs 34503978665 and 34505591306 (session request ndjson):

  • open finishes ok in ~4 s.
  • The wait first tries the AX bridge: ios.snapshot-source.prepare / .acquire end with request canceled at ~5.2 s (the snapshot-source default maxDurationMs is 5 s and its signal aborts the in-flight probe or compile, so the error kind is cancelled).
  • The runner fallback then logs ios_runner_session_reuse → ios_runner_connect ok in 7–59 ms → ios_runner_readiness_preflight with timeoutMs: 1000 → a second ios_runner_connect that stalls 4.5–5.1 s → the wait's own deadline cancels it, captures: 1, readableCaptures: 0, and the wait fails at 10.4 s.
  • A green run without any of the suspected changes shows the same second connect at 11, 12, 18, 19, 75, 203, 1073 and 2323 ms. Same shape, same budgets (preflight 1000 ms, connect ~45 000 ms, command ~44 9xx ms), just inside the margin. A later plain snapshot in the failing session succeeds in ~6 s: the runner is alive, it was busy.

Reading: the 1 s readiness preflight abandons a runner command that the runner keeps executing for ~5 s; the next connect waits behind it; a 10 s wait that already spent 5 s on the bridge has nothing left. Cold hosts make the runner's first commands slow enough to cross the line.

Not the cause

#2473 was suspected and reverted in #2474; the revert was withdrawn after a local reproduction on a booted simulator passed 4/4 on both f4c8f3ddda and its parent with identical request logs, and after the CI budgets were found byte-identical between red and green runs. Two of the correlated failures were a Swift-only XCTest step on a different Xcode image (ALERT_DEADLINE_EXCEEDED), unrelated.

Required behavior

  • A readiness preflight that times out must not leave the runner executing an abandoned command that blocks the next connect, or the next connect must not wait behind it. Options: make the preflight command idempotent and cheap enough that 1 s is honest on a cold host, cancel it on the runner side when the client abandons it, or size the preflight from the request's remaining budget instead of a fixed 1 s.
  • A wait whose snapshot-source attempt consumed ~5 s must hand the runner fallback a real remainder, and the fallback must not spend it on a connect that cannot answer.
  • Evidence to close: the fixture smoke and the settings replay green on a cold runner across several consecutive main runs, and a unit test that plants a 5 s busy runner behind a 1 s preflight and shows the wait still answers within its budget.

Non-goals

Widening the wait's budget in the smoke tests; retrying the smoke harder.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingneeds-triageNew or unreviewed issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions