Skip to content

feat(core): record working directory and completion time on tool calls - #356

Merged
Coly010 merged 7 commits into
mainfrom
columferry/core-tool-call-cwd
Oct 7, 2026
Merged

Coly010 merged 7 commits into
mainfrom
columferry/core-tool-call-cwd

Conversation

@Coly010

@Coly010 Coly010 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Current Behavior

Scorers can't tell which directory an agent ran a tool call in, or when it finished. ToolCallRecord has no cwd, so a CLI eval that runs supabase start from client-a and client-b can't attribute each call to a project. adaptTranscript also drops the result time that #351 already stamps on each tool result, so evals/cli/lib/detours.ts reads both fields through defensive unknown casts.

Expected Behavior

A small change on top of the rollout enrichment #351 added (no new rollout parsing, no sessionLog plumbing):

  • Codex: enrichFromRollout now pairs tool calls with rollout items per call instead of all-or-nothing. Each call takes the next rollout item (in order) that matches its command or MCP tool and sets event.tool.cwd from the item's cwd (a file:// URL is converted to a path). Extra rollout items the --json stream omits, like a killed long-running command, are skipped, and a call with no match stays unenriched without affecting the others. This also improves feat: align agent traces with Braintrust's span spec #351's timestamp enrichment, since the same pairing drives it. File-change and web-search items now only pair with calls of the same kind, so scanning past an extra item can't mispair them.
  • Codex: a command's completion time (resultTs) comes from its rollout item_completed event instead of the function_call_output, which for a yielded long-running command marks the yield rather than the exit (the output arrives early and write_stdin polls the rest). It falls back to the output time when item_completed has no timestamp.
  • OpenCode: bash's workdir is normalized to cwd through the shared arg extractor.
  • adaptTranscript copies cwd onto ToolCallRecord, and also the result time as ToolCallRecord.resultTs. It is the same value and name as TranscriptPart.resultTs, so no new endedAt field.
  • evals/cli/lib/detours.ts reads record.cwd and record.resultTs directly. The casts and the cast-based test helper are gone.

Replay check

I replayed the real enrichFromRollout against the Codex runs in the three most recent main CI raw-results artifacts (279 runs; calls rebuilt from each result.json, rollouts from session-archive.tar.gz), before and after the per-call pairing. Counts are command calls with a cwd and timestamp.

all commands enriched some enriched none enriched
all-or-nothing (#351) 249 0 30
per-call pairing 274 5 0
  • The 249 runs that already paired are unchanged: their enriched events are identical before and after.
  • The 30 runs that were fully unenriched are all build-functions-006 / build-functions-007. 25 had one extra rollout CommandExecution the stream never emitted (failed, exit -1: a long-running supabase functions serve killed after ~50s); they are now fully enriched.
  • 5 runs (all build-functions-007) are still partially enriched: 13 of their command calls stay unpaired, 12 of them apply_patch / git apply / patch / cat > heredoc commands whose rollout argv differs from the stream's command text, and 1 docker exec ... sh -c '...' command that also doesn't match its rollout argv. Everything around them is enriched. Matching those heredocs would need a looser comparison, which this PR doesn't attempt.

Test plan

  • pnpm format:check, pnpm typecheck
  • pnpm --filter @supabase-evals/core test
  • pnpm --filter @supabase-evals/framework exec vitest run --root ../.. evals/cli
  • pnpm --filter @supabase-evals/framework test:cli-lib
  • New tests: per-call Codex pairing (extra rollout item, one mismatched command, fewer rollout items, order safety); Codex paired item with a path / file:// / no cwd; OpenCode workdir to cwd; adaptTranscript copies cwd and resultTs; extractCommandEntries with typed records.
  • packages/core/src/agents/codex/parser.test.ts: the MCP transcript assertion now includes id: 'item_7'. It was failing on main since feat: align agent traces with Braintrust's span spec #351.

🤖 Generated with Claude Code

@Coly010
Coly010 requested a review from a team as a code owner October 1, 2026 12:08
@vercel

vercel Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
evals Ignored Ignored Preview Oct 7, 2026 10:50am UTC

Request Review

@Rodriguespn Rodriguespn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now that #351 is merged, most of this PR already lives on main: enrichEvents reads the rollout, and enrichFromRollout pairs each rollout item with its tool call (sameCall / splitShellWords, plus redacted-secret matching). So sessionLog / ParseContext.sessionLog and the separate rollout parsing can go.

I think #356 can be compacted to:

One thing to check: #351 only enriches tool calls when all of them match, whereas yours matched each command on its own.
Could you replay your 30 CI rollouts to see whether any run loses all its cwds?

Conflicts in the codex parser/runner, engine, and agent types are resolved
by taking main's #351 rollout enrichment; this PR's sessionLog plumbing and
rollout parsing are dropped in favour of it.
Set each Codex tool call's cwd from the rollout item #351 already pairs it
with, copy it and the result time onto ToolCallRecord, and read both
directly in the CLI detour helpers.
@Coly010

Coly010 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Thanks Pedro, that's a much smaller PR now 🙂 merged main and compacted it to your shape:

  1. codex cwd is set from the rollout item inside feat: align agent traces with Braintrust's span spec #351's pairing (kept the file:// helper), and sessionLog plus the separate rollout parsing are gone
  2. opencode workdir → cwd is unchanged
  3. instead of adding endedAt, adaptTranscript copies feat: align agent traces with Braintrust's span spec #351's resultTs onto ToolCallRecord next to cwd, so there's one name for the completion time across agents
  4. the defensive casts in evals/cli/lib/detours.ts are gone, and it reads the typed fields directly. I skipped the braintrust span metadata for now since TranscriptPart has no cwd, so it'd need a schema change rather than a one-liner, happy to follow up if you want it

On the replay: the original 30 rollouts had expired, so I replayed 279 runs from main's three latest CI artifacts instead. All 45 build-database-002 runs pair fully. 30 runs lose every cwd to the all-or-nothing match, all of them build-functions-006/007: 25 have one extra killed supabase functions serve in the rollout that --json never emits, and 5 are heredoc apply_patch/git apply argv mismatches. Matching in order and letting a call skip unmatched rollout items recovers the 25. That'd change #351's timestamp pairing too though, so it feels like a follow-up rather than something to sneak in here. Thoughts?

Also fixed the codex parser test that was red on main after #351 (id: 'item_7'), happy to pull that out if you'd rather fix it separately

@Coly010
Coly010 requested a review from Rodriguespn October 6, 2026 09:31
@Coly010 Coly010 changed the title feat(core): record the working directory of agent tool calls feat(core): record working directory and completion time on tool calls Oct 6, 2026
@Coly010

Coly010 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Quick update - changed my mind on the follow-up, since the PR's title doesn't really hold if a run can lose every cwd. Pairing is per call now: each command scans forward for its rollout item and skips extras (like the killed functions serve), and an unmatched call is left alone without throwing off the rest. Replay: runs losing all enrichment went from 30 to 0, the 249 that already paired are byte-identical, and 13 heredoc/docker exec calls across 5 runs stay unpaired with their neighbours still enriched. One knock-on: sameCall now requires a file_change/web_search call for those item types, so a command can't pair with a file change while scanning forward

@Rodriguespn Rodriguespn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replayed 186 codex runs from the last two main refreshes: 0 runs lose all cwds, 6 unpaired calls, all apply_patch/heredoc/docker exec curl with a redacted secret (the redactor eats a "), none are supabase invocations. Good to go from my side 🚢

@Coly010
Coly010 merged commit acb7fb4 into main Oct 7, 2026
6 checks passed
Coly010 added a commit that referenced this pull request Oct 7, 2026
… landed

Read resultTs in the fleet test fixtures, find service stacks recreated under another name, and update the README now that #356 is merged.
Coly010 added a commit that referenced this pull request Oct 7, 2026
## Current Behavior

#356 records the working directory of each tool call on
`ToolCallRecord`, but it stops there. The Braintrust upload builds tool
spans from transcript parts, which carry no `cwd`, so a trace can't show
which project a `supabase start` or `supabase stop` ran in. This was
requested in the #356 review.

## Expected Behavior

Tool-call transcript parts carry an optional `cwd`, copied from the
event's `tool.cwd` in `adaptTranscript`. The upload's transcript schema
accepts it and tool spans log it as `metadata.cwd` next to `tool_name`.
The key is left off when the agent didn't record a directory. Results
written before this change still parse.

## Test plan

- [x] `adaptTranscript` sets `cwd` on the transcript part only when the
event has `tool.cwd`
- [x] Transcript schema parses parts with and without `cwd`
- [x] `logTranscript` puts `cwd` in tool span metadata when present and
omits it otherwise
- [x] `pnpm format:check`, `pnpm typecheck`
- [x] core, framework, sandbox, web tests and `test:cli-lib`
- [ ] `pnpm check` end to end (the judge smoke step needs a provider
key)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Coly010 added a commit that referenced this pull request Oct 7, 2026
…LI-2402) (#316)

## Current Behavior

Steady-state agent work is managing a fleet of local stacks, not
starting one. Being able to find, restart and remove the right stack is
part of the launch scope for parallel stacks, and nothing measures it
separately from the startup scenarios.

## Expected Behavior

Adds `evals/cli/resolve-database-002-stale-stack-cleanup` (`interface:
cli`, `stage: resolve`, `projectRunning: false`, `needsDocker: false`).
The agent:

1. Stands up local Supabase for `checkout-service`, `payments-api` and
an old `legacy-import` prototype, each with a `service_marker` row
naming it.
2. Tears `legacy-import` down completely.
3. Restarts `checkout-service`.
4. Leaves `payments-api` exactly as it is.

The prompt never names a command or Docker.

Structured like `build-database-002`: `EVAL.ts` composes the checks and
owns both `judge()` calls; `services.ts`, `fleet.ts` and `metrics.ts`
hold the deterministic checks, each with its own tests. This PR also
adds one shared module, and carries one shared-lib fix:

- `lib/stack-list.ts`: parses `stack list` output (a `{ stacks }`
envelope or a top-level array). Only an unknown-subcommand error marks
the listing unsupported; any other unreadable output fails closed.
- `lib/stack.ts` (named managed stacks): `resolveStack` now finds named
managed stacks per project (from `stack list`, filtered by project root)
and reads `runtime` as either a string or `{ kind }`. Agents in CI fell
back to `--stack native`, `--stack checkout-service-recovered` and
similar names that the scorer couldn't resolve. These `lib/stack.ts`,
`lib/stack.test.ts` and `lib/README.md` changes are identical to #314's
(`6f7041d`), so whichever PR merges first, the other merges cleanly. On
top of it, this eval retries a service that doesn't resolve by its own
name without the name, so a stack recreated under another name in that
service's directory still counts. Each distinct probe command runs once
per scoring pass, so that retry doesn't repeat the probes or notes of
the first lookup.

It uses the shared `lib/cli-invocations.ts` (#354, merged) to find the
`supabase` commands the agent actually ran (argv-parsed, skipping echoed
or passive text). **Attribution is directory-first, except
`--project-id`**: `--project-id <name>` is authoritative, so `cd
checkout-service && supabase stop --project-id payments-api` stops
payments-api. Otherwise an invocation targets the service whose
directory it ran in (`--workdir`/`SUPABASE_WORKDIR` basename, else an
earlier `cd`/`pushd`, else the tool call's `cwd`), and only when that
directory is unknown or isn't a service directory does `--stack` decide,
so `cd legacy-import && supabase stack destroy --stack payments-api`
targets legacy-import. The shared lib (`lib/cli-invocations.ts`,
`lib/detours.ts`, `lib/metrics.ts` and their tests) is byte-identical to
#314's, synced in `e29c3ae`, so whichever PR merges first, the other
merges cleanly.

The README explains what the eval measures and what each experiment is
expected to show.

| Check | Proves |
|---|---|
| checkout-service and payments-api projects exist | found per service,
so deleting `legacy-import`'s directory doesn't fail the survivors; by
`supabase/config.toml`, else by the CLI's `stack list` (see below) |
| checkout-service / payments-api stack is running | each resolves
(managed-named in the project dir → managed-named at root → managed →
legacy) and answers `select 1` |
| legacy-import stack is gone | decided by end state: not listed
(default and relocated homes), not resolvable, nothing on its configured
db port, and no leftover labelled containers once its directory is gone,
**and** a `supabase` start targeted it that didn't fail (so it can't
pass vacuously). The teardown command is reported, not required |
| checkout-service was restarted | its Postgres postmaster started after
setup completed (state); without timing, a `stack restart` (or stop +
start) targeting it in the change phase, not failed (commands) |
| payments-api left untouched | no `db reset` targeted it in the change
phase, and its postmaster started before setup completed (state) or,
without timing, no stop/restart/destroy targeted it (commands); it still
holds its own marker row |
| surviving stacks kept their data | `service_marker` isolation, both
directions, case-insensitive |
| no container-runtime detours | LLM judge over the complete executed
commands (shared rubric; stopping Supabase's own containers with `docker
stop` isn't a detour, it just doesn't count as a teardown) |
| metrics | per-service backend/runtime, ports, postmaster start times,
`setupCompletedAt`, which evidence decided each check, `legacyTeardown`
(`succeeded`/`failed`/`none`), `cliDetours` — reported, never asserted |
| final report is truthful about the fleet | LLM judge given harness
ground truth, built from the same decisions as the checks and naming the
evidence that decided them |

Decisions worth a look:

- **"Completely" means absent from everything the CLI reports.** A
managed `stack stop` that leaves the stack listed fails; `stack destroy`
or legacy `supabase stop` passes.
- **Evidence model: end state first, commands as fallback.** The setup
point is the start that completes setup (every service has had a start
that didn't fail and that names it); its time is that tool call's
recorded completion time. When both that time and a survivor's
`pg_postmaster_start_time()` are known, and the survivor wasn't stopped
or restarted in the same tool call as setup (one completion time can't
order them; that falls back to commands, saying so), the postmaster
decides: started more than 1s (`CLOCK_TOLERANCE_MS`) after setup means
restarted, otherwise not, and commands aren't consulted. So a restart
the parser can't attribute still counts for checkout, and a "successful"
`stack restart` that left Postgres running doesn't. Without both, the
command rules below decide. The setup time is the call's completion time
(`resultTs`), never an issue time. #356 has merged, so
`ToolCallRecord.resultTs` and `ToolCallRecord.cwd` are typed core fields
and timing evidence is live wherever an agent records them (Codex's
rollout does); runs without them fall back to command evidence. Setup is
anchored on the latest completion time among each service's first such
start, since parallel starts finish out of order. A shell-loop start
(`cd "$s"`) is start evidence for every service but never completes
setup or becomes the anchor while every service has a start that names
it; only when some service has none does setup complete on the first
point all three are covered, anchored on the latest loop start, and
state evidence is skipped (restart and untouched fall back to commands,
saying so). A completion time that belongs to no one service can't order
that service's postmaster.
- **`db reset` is a database restart.** A local reset recreates
(Postgres 15+) or restarts (Postgres 14) the database container, so it
moves the postmaster start time. In the change phase it always counts
against payments-api (even if the command failed) and never counts as
restarting checkout-service on its own: under state evidence checkout
also needs a `stack restart` or a stop then start that didn't fail. The
change phase begins at the first stop, restart or destroy of any service
after setup; a `db reset` before it is seeding, part of setup, and moves
a service's reference time to its completion, so seeding every service
via `db reset` neither fails payments-api nor reads as a checkout
restart. A seeding reset that would move the reference time and shares a
tool call with that service's change-phase stop, restart or destroy
can't be ordered by one completion time, so that falls back to commands,
saying so.
- **Every fleet check names its evidence.** Notes start with `state:`,
`commands:` or `unavailable:` (legacy-import lists each gate:
`listing:`, `resolution:`, `db port:`, `containers:`, `commands:`),
citing commands as `cmd #<n> "<argv>"` (1-based, truncated to 120 chars)
and ISO times. The running checks say whether the stack resolved.
- **legacy-import "gone" is decided by end state; the teardown command
is reported, not required.** In CI, `supabase stop --no-backup` exited 1
with `StopVolumePruneError`/`LegacyStopVolumePruneError`
([CLI-2637](https://linear.app/supabase/issue/CLI-2637)) on stacks that
were really removed, which would have failed the check; in run
37625367762 all 3 such teardowns passed by state. A successful start of
`legacy-import` is still required (start evidence keeps loop starts and
`--stack-id` mapping), so a stack that never ran can't pass. The notes
always say how teardown went: `teardown cmd #N "…" (succeeded)`,
`(failed: StopVolumePruneError, see CLI-2637)` plus `(product gap
CLI-2637)`, or `no teardown command found` (e.g. stopped through
`docker` directly). The judge's ground truth states the same, and
`metrics.legacyTeardown` records it. `psql -c "delete …"` or `rm -rf
legacy-import` still leave the stack resolving or listed, so the state
gates fail them.
- **Latest start decides a CLI-version swap.** A survivor whose latest
start that didn't fail (by completion time when recorded, else command
order) ran through `npx`/`bunx`/`pnpm dlx` of another explicit version
fails `… stack is running`; one whose latest non-failed start used the
installed CLI passes, and a global `npm i -g supabase@X` applies to
later invocations only when its tool call succeeded (a failed install
leaves the installed CLI in place; a global uninstall clears it). The
installed version is the marker's `cliVersion`, else `/usr/bin/supabase
--version`, else PATH; `metrics.cliVersion` is that staged version and
`cliVersionAfterRun` the PATH one when it differs. Dist-tag runners
(`supabase@beta`) can't be checked offline and are reported in
`metrics.cliRunnerUnverified`.
- **Services are found by `config.toml`, else by the CLI's `stack
list`.** CI agents often created stacks without `supabase init` (`mkdir
-p payments-api && cd payments-api && supabase stack start --runtime
native --stack payments-api`) or ran named stacks from the sandbox root,
sometimes under a relocated `HOME`/`SUPABASE_HOME`, so the `config.toml`
search found nothing. For each service it missed, the scorer reads
`stack list` under the default home and each relocated home the agent
started the service with, and takes entries whose `name` is the service
(reachable preferred). One `project_root` is the project directory
(`(from stack list)`); the sandbox root itself counts as a named stack
started there (`named stack at sandbox root`, resolved by the existing
named lookup); two or more distinct roots are `ambiguous (stack list:
…)` and fail. When a `config.toml` exists the directory is unchanged,
but if resolving the stack there fails (an agent broke its own
`config.toml`, or rooted the stack elsewhere), the one `stack list`
entry named exactly for the service gives the directory to resolve in,
under the default and relocated homes; ambiguous roots stay unresolved.
`lib/projects.ts` is untouched.
- **Failed commands are never evidence, except starts that reported
ready.** A tool call that exited non-zero, or whose output holds a CLI
error (`"_tag":"Error"`, `Unknown subcommand`, `UnknownSubcommand`),
doesn't count as a start, teardown or restart. A call with no recorded
outcome counts as succeeded. Top-level `supabase restart` never counts,
since neither CLI has it. **Start-marker exception:** when a failed
call's output has no CLI error envelope, each `[task] done: Stack is
ready.` or `Started supabase local development setup` it printed credits
one start invocation of that call (in order) as not failed. A start that
came up isn't discarded because a later command in the same call failed
(`… > out && jq …` exiting 127, or a `set -e` seed that hit a TLS
error). Other verbs keep the all-failed rule.
- **Change phase.** Command evidence for restarts and touches only
counts after the setup point, so setup-phase stop/start retries (e.g.
clearing a port clash) are neither a restart nor a touch. If that point
never comes, the restart and untouched checks fail with an
`unavailable:` note.
- **Marker rows match their service name as a whole word**
(case-insensitive), the same rule as the isolation check, so `marker for
payments-api` is payments-api's marker and `payments-api-archive` isn't.
- **Leftover legacy containers.** Once `legacy-import`'s directory is
gone, the scorer checks `docker ps` for containers labelled
`com.supabase.cli.project=legacy-import` or named
`supabase_*_legacy-import`; skipped when `docker` is unreachable, since
then nothing can be running. The scorer never reads which experiment it
runs under.
- **`stack list` fails closed.** Only an unknown-subcommand error skips
the listing half; any other unreadable output fails `legacy-import stack
is gone`.
- **Truthful judge.** When nothing was started, a single clear statement
covering all three services is enough, so an honest blanket report isn't
failed as vague. When the database start time says checkout wasn't
restarted but a restart command after setup didn't fail, the ground
truth states both, and the rubric treats an agent that accurately
reports the CLI-reported successful restart as truthful; the check
itself still follows state evidence.
- **`--stack-id` is mapped to a service** from the `"id"` a successful
start printed (when that start targeted one service) and from any `stack
list` output the agent printed; an unknown id is attributed to no
service.
- **Shell loops** (`for s in …; do (cd "$s" && supabase start); done`)
count as start evidence for every service (see the anchor rule above).
Loop teardowns and restarts count for none, since their target can't be
attributed.
- **Known limitations** (in the README):
- Failure is known per tool call, so one failing command in a compound
call marks all of its invocations failed. Starts are the exception: a
failed call's start counts when the output printed a ready marker and
holds no CLI error, one start per marker, in order.
  - `cd` doesn't carry across separate tool calls.
- It's unverified whether a native `stack restart` restarts Postgres; if
it doesn't, state evidence on a native stack fails `checkout-service was
restarted` after a restart that succeeded (the truthful judge is told
about both).
- The container probe assumes legacy-import's project id is its
directory name.

The CLI-2402 issue describes the CLI side as blocked. It no longer is:
`stack list/start/restart/destroy/status --stack` exist in beta behind
`SUPABASE_EXPERIMENTAL_STACK=1`. The harness can't hand an agent
pre-running named stacks, so the prompt is two-phase: build the fleet,
then manage it.

## Results

CI run 37625367762 scored 6/15.

- Three `legacy-import` teardowns failed with `StopVolumePruneError`
([CLI-2637](https://linear.app/supabase/issue/CLI-2637)) and passed by
state.
- Six agents never started a stack.
- pinned r3 swapped the CLI to `npx` 2.120.0 (policy fail; it then hit
CLI-2638).
- stable r2 used `stack stop`, which leaves the stack listed (correct
fail).
- nodaemon r2 broke its own `config.toml` (a duplicate `[experimental]`
table gives `CliConfigParseError`) and rooted the stack at
`checkout-service/runtime`, so the scorer said it wasn't running.
Resolution now falls back to the `stack list` entry for the service.
- beta r1 was a scorer bug, now fixed. It ran `supabase start --workdir
"$PWD"` once per service directory; `$PWD` wasn't expanded, so the first
start read as a loop start, anchored setup too early, and payments' own
start read as after setup (`payments-api left untouched` failed, with no
teardown or restart attributed either). `$PWD` is now expanded against
each invocation's cwd (shared `lib/cli-invocations.ts`, identical to
#314), unresolved starts no longer anchor setup, and the replay passes
all fleet checks.

#355 (the sandbox login-shell `PATH` fix) is merged, so `nodaemon` is
now real; earlier `nodaemon` results were affected by that bug. The
per-call working-directory gap is fixed in #356, which is now merged, so
per-call `cwd` attribution and timing evidence are in effect. Some
Docker-arm agents decline to start because `docker ps` shows the
sandbox's own container; that's an environment artifact.

Known product gaps seen in CI:

- [CLI-2637](https://linear.app/supabase/issue/CLI-2637): `stop
--no-backup` exits 1 with `StopVolumePruneError` on Docker clients older
than 23.0, even though the containers are removed
- [CLI-2638](https://linear.app/supabase/issue/CLI-2638): the managed
Docker runtime fails to bind-mount when the daemon doesn't share the
CLI's filesystem (the sandbox's sibling daemon)
- [CLI-2639](https://linear.app/supabase/issue/CLI-2639): `stack
destroy` can't remove a stack whose launch failed
- [CLI-2640](https://linear.app/supabase/issue/CLI-2640):
Docker-unavailable errors don't mention `--runtime native`, so some
Docker-less agents give up

## Related Issue(s)

Refs https://linear.app/supabase/issue/CLI-2402

## Test plan

- [x] `vitest run --root ../.. evals/cli` (counterexamples include psql
delete, `rm -rf`, never-started stop, a failed `stop --no-backup` with
`StopVolumePruneError` and clear state (passes with the CLI-2637 note),
no teardown command with clear state (passes), a started stack still
listed (fails), services found via `stack list` (default home, relocated
home, sandbox root, ambiguous roots), loop starts, echoed restart,
`stack restart --stack payments-api`, directory removed after destroy,
failed `stack restart`, top-level `supabase restart`, setup-phase
stop/start retries, failed legacy teardown after `rm -rf`, leftover
labelled containers, unreadable `stack list` output, `db reset --workdir
payments-api`, a checkout `db reset` that doesn't count as a restart,
seeding every service via `db reset` before any change, a seeding `db
reset` sharing a call with the restart (commands decide) versus in an
earlier call (state decides), `stop --project-id payments-api` run from
another service's directory, a failed versus successful global install,
a loop start followed by per-service starts, only loop starts, the beta
r1 replay, a broken `config.toml` found via `stack list`, a `marker for
payments-api` row, failed starts in the swap rule, `--stack checkout`
and `--stack checkout-service-recovered` started in the service
directory, dist-tag runners not counted as overrides, no repeated probe
commands in the renamed-stack lookup; state evidence: checkout restarted
with no restart command, "successful" restart with an older postmaster,
payments restarted by state, payments stop command overruled by state,
`db reset` under state, fallback to commands without timing, ignoring a
tool call issue time, setup and restart/touch in one call, ground truth
stating both state and a successful restart command)
- [x] `apps/framework` `tsc --noEmit`
- [x] `pnpm eval:dry -- --suite cli --experiment-suite cli` → planned
under all 5 cli experiments
- [x] `pnpm format:check`
- [x] `pnpm check`, all steps except the judge smoke step, which needs a
provider API key locally
- [x] CI refresh via `run-evals-changed`: inspect one passing run and
each distinct failure shape
- [x] Merged `origin/main` (#356 included); `pnpm check` and the rest
re-run after the merge
- [x] Eval results refresh on this push (`run-evals-changed`): confirm
timing evidence and per-call `cwd` attribution show up, incl.
`nodaemon`, and that renamed stacks resolve

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Coly010 added a commit that referenced this pull request Oct 7, 2026
…399) (#314)

## Current Behavior

Port collisions between concurrent local Supabase projects are a
long-running pain point ([discussion
#5968](https://github.com/orgs/supabase/discussions/5968)). The CLI
historically assumed one machine means one fixed port set. Running
parallel stacks is the flagship behaviour that work is meant to fix, and
nothing measures whether an agent can actually drive it.

## Expected Behavior

Adds `evals/cli/build-database-003-parallel-projects` (`interface: cli`,
`projectRunning: false`, `needsDocker: false`). The agent stands up two
independent local Supabase projects, `client-a` and `client-b`,
concurrently in one sandbox. Each project gets a `clients` table holding
a row naming that client, and the agent reports which API port each one
landed on. The prompt never names a command, a port or Docker.

Structured like `build-database-002`: `EVAL.ts` composes the checks and
owns both `ctx.judge()` calls; `projects.ts`, `stacks.ts`, `report.ts`
and `metrics.ts` hold the deterministic checks, each with its own tests.
Shared probing (stack resolution, project discovery, marker reads,
detour policy) comes from `evals/cli/lib`. The README explains what the
eval measures and what each experiment is expected to show.

| Check | Proves |
|---|---|
| two projects initialised | both project dirs found via
`*/supabase/config.toml`, exact basename preferred (`client-a-old`
doesn't make `client-a` ambiguous) |
| both stacks reach ready | each project's stack resolves from its own
directory and answers `select 1` |
| stacks are on distinct ports | two independent stacks, not one reused
or clobbered |
| each project holds only its own marker row | both directions,
case-insensitive: `client-a`'s db has `client-a` and not `client-b`, and
vice versa |
| each clients table holds exactly one row | the prompt asks for a
single row per client; duplicates or an empty table fail |
| reported api ports match the running stacks | each project's real API
port appears in the final report as a whole number |
| no container-runtime detours | LLM judge over the complete executed
commands only (shared rubric with 002) |
| metrics | per-project backend/runtime, ports, time to ready,
`cliDetours` (regex diagnostic) — reported, never asserted |
| final report is truthful about both projects | LLM judge given harness
ground truth plus the transcript with tool outputs; fails on a misstated
outcome, rows or ports, or an omitted/swapped API port, not on rows the
report didn't recite |

Known-bad cases covered by unit tests:

- a port that only appears inside a longer number
- rows swapped between the two databases
- one database holding both rows
- `cd client-b` resolving to `client-a`'s stack
- a stack that resolves but reports no API URL, which fails with a clear
note instead of falling back to `config.toml`
- a `clients` table with three rows, or none

The CLI-2399 issue describes this eval as blocked on multi-stack session
support. That no longer applies: with `projectRunning: false` the
harness starts zero stacks and only installs the CLI, so running two
stacks is entirely the agent's job, which is the point of the scenario.

**Expected results**: every experiment should pass. `absent` and
`nodaemon` pass through `--runtime native` (pinned to beta, since the
native managed stack only exists there); the Docker arms need the agent
to move one project off the default ports unless the CLI allocates them.

**Observed in CI** (snapshot: run 37638547341, 2026-10-07, head
76e9c61): 12/15 runs pass. Results drift between runs.

| experiment | passed |
|---|---|
| `absent` | 3/3 |
| `nodaemon` | 3/3 |
| `beta` | 3/3 |
| `pinned` | 2/3 |
| `stable` | 1/3 |

- `pinned` r3 fails the version-swap rule: the agent ran `npx --yes
supabase@2.120.0`.
- `stable` r2 and r3 fail `both stacks reach ready` with no start
attempted for either project.

**Policies**

- Invocations are attributed to a project by the directory they ran in
(`cwd`, `--workdir`, `SUPABASE_WORKDIR`, `env -C`) before their
`--stack` name.
- Relocated home passes, visibly: when the default home can't resolve a
stack, the scorer retries under each `SUPABASE_HOME`/`HOME` the agent
set on a start of that project (prefix, `env`, or an earlier `export` in
the same command; a relative value resolves where the CLI runs). The
`relocated home` note and the `relocatedHome` metric keep it visible.
- Version swap fails, visibly: each project is judged by its latest
start. If it used an `npx`/`bunx`/`npm exec`/`pnpm dlx`/`yarn dlx`
runner or followed a global reinstall of another explicit version
(compared against the staged CLI version, cleared by a global uninstall,
and ignored when the PATH version after the run still equals the staged
one), `both stacks reach ready` fails (`cliOverride` metric); a latest
start through the installed CLI passes. The truthful-report judge's
override note is added only to projects that were started that way.
Dist-tag runners (`@latest`, `@beta`) can't be verified offline and are
recorded in `cliRunnerUnverified` instead.

## Related Issue(s)

Refs https://linear.app/supabase/issue/CLI-2399

## Test plan

- [x] `vitest run --root ../.. evals/cli`, `test:cli-lib`, and `pnpm
typecheck`, re-verified against merged `main` (including #356)
- [x] `apps/framework` `tsc --noEmit`
- [x] `pnpm eval:dry -- --suite cli --experiment-suite cli` → 003
planned under all 5 cli experiments
- [x] `pnpm format:check`
- [x] `pnpm check`, all steps except the judge smoke step, which needs a
provider API key locally
- [x] CI refresh via `run-evals-changed`: inspected passing runs and
each distinct failure shape
- [x] Results refreshed via `run-evals-changed` (run 37638547341)

Also carries #316's final commit (post-run install reconciliation wiring
for stale-stack-cleanup), which missed #316's merge, so the shared lib
and both evals stay consistent on main.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants