Skip to content

feat(evals): add build-cli-004-worktree-stacks (CLI-2400) - #295

Closed
kanadgupta wants to merge 7 commits into
mainfrom
kanad-claude/cli-2400-worktree-stacks
Closed

kanadgupta wants to merge 7 commits into
mainfrom
kanad-claude/cli-2400-worktree-stacks

Conversation

@kanadgupta

Copy link
Copy Markdown
Member

Tracks CLI-2400 under the Slim CLI evals RFC.

Note

Stacked on #281. This branch is Colum's branch plus origin/main merged in plus the commits below. The diff will shrink to the last two commits once #281 lands.

The problem it's solving

The headline agent workflow for the Slim CLI launch is one human plus N coding agents, each working in its own git worktree, each needing its own isolated local Supabase stack. Nothing in this repo measured whether an agent can actually get there: three worktrees, three stacks, zero leakage.

What the PR adds

One new regression eval, evals/regression/build-cli-004-worktree-stacks/. The agent is asked to set a repo up with worktrees feature-a, feature-b, feature-c, start a local stack in each, and add a different table (widgets / gadgets / gizmos) plus one sample row per worktree. No seed data: the agent builds everything, including the worktrees.

A scorer that grades the end state, never the method. Because the harness's built-in ctx.query / stackStatus resolve a single stack from the workspace root, every check addresses stacks itself by cd-ing into each worktree:

  • three real git worktrees on distinct branches (three plain directories don't count);
  • one live stack per worktree with three distinct database endpoints, which is what catches aliasing or reuse;
  • per-table schema isolation via to_regclass against all three stacks;
  • at least one row per table in its home stack;
  • each table created by a migration file in its worktree;
  • an always-passing metrics check reporting fleet wall-clock across the three stacks, CLI version and channel, per-stack backend and runtime, and the number of supabase start invocations.

The scorer asks the managed backend first (SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env --output-format json) and falls back to the legacy one, so it works whichever path the agent took and however it enabled the flag. Same environment-agnostic shape as #281: it never branches on which experiment arm it runs under.

Two experiment edits. The eval id is added to the Docker-less allowlists in -cli-nodaemon and -cli-absent. The pinned Codex Luna experiment and the -cli-stable / -cli-beta arms already pick it up through their suite and interface rules, so the eval runs under all five. No skipEval anywhere.

One housekeeping commit. main now discovers evals as evals/<suite>/<id>/ and parses every top-level folder as a suite name, so #281's flat evals/build-database-002-stack-lifecycle/ breaks discovery once main is merged in. This branch moves it into evals/regression/ and drops the now-ignored suite: frontmatter line. That commit becomes redundant when #281 rebases.

Deliberate choices

  • Nothing tells the agent about the managed stack or its feature flag. Worktree isolation only exists on the managed stack (beta channel, gated behind experimental.stack). Whether agents can discover that is part of what's measured. Expect every arm to fail until they do or the CLI flips the default.
  • services: is omitted. With it set, the sandbox shim appends legacy container names (-x gotrue,kong,...) to supabase start, and the managed backend rejects those outright (Unknown stack capabilities in --exclude). This also affects feat(evals): add build-database-002-stack-lifecycle CLI eval (CLI-2398) #281's eval, which sets services: [].
  • Known footgun left in. supabase init still pins ports in config.toml. The managed stack treats them as exact intents, so the second worktree's start fails with Persisted database port is unavailable until the ports are removed and become automatic. That is exactly the "no port surgery" promise the launch makes.

Verification

Verified by hand against CLI 2.118.0-beta.37 on the native runtime: three sibling worktrees get three distinct stack identities with no config edits, and a table created in one is absent from the others. The managed backend rejects supabase status -o json, the DB password is random per stack, and excluding auth on this beta fails start with Persisted API gateway material is incomplete.

  • Scorer unit tests (scoring.test.ts, 23 cases; fixtures captured from real beta CLI output)
  • pnpm typecheck, Biome
  • Harness dry run plans the eval under the pinned, beta, and absent arms
  • Real scorer executed against three hand-built worktree stacks on the beta CLI: all checks pass; stopping a stack or leaking a table fails the right checks with readable notes
  • Agent-backed runs. I have no Anthropic or OpenAI API keys locally, so results come from CI via the run-evals-changed label on this PR.

🤖 Generated with Claude Code

Coly010 and others added 6 commits September 11, 2026 16:30
…periments

Three regression evals for CLI-2398 share one prompt (init a project, start
the local stack, add a seeded `notes` table) and one scorer, differing only
in the sandbox's Docker state declared beside PROMPT.md in
sandbox-environment.json: 002 is the Docker-available control, 003 has the
docker client but no usable daemon, 004 has no docker binary at all. The
scorer checks stack readiness, the seeded rows, zero container-runtime
detours (install/start/escalation attempts), records cliVersion,
resolvedRuntime, timeToReadyMs and raw-socket probes as metrics, and asks an
LLM judge whether the agent named the real blocker when it failed. The two
Docker-less arms fail today by design; they track the gap the CLI's native
managed stack is meant to close.

Two experiments clone codex-gpt-5.6-luna but install the latest stable or
latest beta Supabase CLI (resolved lazily from npm dist-tags) via an
experiment-land LocalStackRuntime built from the sandbox package's exports,
so the pinned 2.67.1 baseline, stable and beta can be compared nightly
across every interface: cli regression eval. The Docker-less arms are staged
with DOCKER_HOST pointed at an unbound port plus root-owned PATH shims (the
CI sandbox makes the socket world-writable, so permissions alone cannot
block it); for 003 the docker shim still answers --version so the managed
stack's runtime probe selects Docker as it would on a real host.

The existing Luna experiments skip evals that need a Docker-less sandbox.
Unit tests for the pure helpers run with:
pnpm --filter @supabase-evals/framework exec vitest run --root ../.. experiments/_lib evals/build-database-002-stack-lifecycle

Refs: https://linear.app/supabase/issue/CLI-2398/add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle
…-2398)

An eval told in advance which environment it runs in, and graded toward
that environment's expected story, measures conformance to an answer key
rather than adaptation — and the marker-file design forced skipEval patches
onto two unrelated pinned experiments just to suppress meaningless nightly
fails. Collapse to one eval, build-database-002-stack-lifecycle, whose
scorer asserts only environment-agnostic criteria (stack ready however it
got there, seeded rows verified, zero container-runtime detours, a truthful
final report) and reports resolvedRuntime as an observed metric.

The forced environments become experiment variants: -cli-nodaemon (docker
client present, daemon unreachable) and -cli-absent (no docker at all) wrap
dockerAwareLocalStackRuntime with a `docker` option instead of reading an
eval-side sandbox-environment.json, on the beta channel where the managed
stack's Docker-less path lives, scoped to this scenario. The existing Luna
experiments return to their upstream content and run the eval under their
stock sandbox like any other.

Also mirrors the upstream promptAddendum/buildSkillsPrompt API change
(cd5b8d0, 03467fb): CLI agents get an empty addendum via
buildToolSurfaceAddendum, matching the stock runtime.

Refs: https://linear.app/supabase/issue/CLI-2398/add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle
…folder

main now discovers evals as `evals/<suite>/<id>/` (#282) and derives the
suite from the folder name, parsing every top-level directory under
`evals/` as a suite. With main merged into this branch, the flat
`evals/build-database-002-stack-lifecycle/` folder from PR #281 would make
`evalSuiteSchema.parse('build-database-002-stack-lifecycle')` throw and
abort discovery for every run.

- Move the eval to `evals/regression/`, matching the `suite: regression`
  it already declared.
- Drop the `suite:` frontmatter line: the runner ignores it in favor of
  the folder, so keeping it would be a second source of truth.
- Update the vitest run hints and the schema cross-reference comment in
  `experiments/_lib` to the new path.

The experiment allowlists reference the eval by id only, so they are
unaffected. This commit is expected to become redundant once PR #281
rebases onto main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One human plus N coding agents, each in its own git worktree, is the
headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree
needs its own isolated local Supabase stack with zero leakage between
them. Nothing in this repo measured that.

Scenario: the agent sets up a repo with worktrees feature-a/b/c, starts a
local stack in each, and adds a different table (widgets/gadgets/gizmos)
plus one seed row per worktree. No seed data: the agent builds everything,
including the worktrees, so `local/` is not needed and no framework setup
hook is required.

Scorer (end state only, all via ctx.exec because the harness's built-in
ctx.query/stackStatus resolve a single stack from the workspace root):

- three real git worktrees on distinct branches (`git worktree list
  --porcelain`), so three plain directories don't count;
- one live stack per worktree with three distinct host:port database
  endpoints, which is what catches aliasing/reuse;
- per-table schema isolation via to_regclass against all three stacks;
- at least one row per table in its home stack;
- each table created by a migration file in its worktree;
- an always-passing metrics check reporting fleet wall-clock across the
  three stacks, CLI version/channel, per-stack backend and runtime, and
  the number of `supabase start` invocations.

Stack resolution asks the managed backend first
(SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back
to the legacy one (SUPABASE_EXPERIMENTAL_STACK=0 supabase status -o json).
Verified against CLI 2.118.0-beta.37: the managed backend rejects the
legacy -o flag outright, the DB password is random per stack, and the
legacy backend cannot see managed stacks, so neither probe alone works.

Deliberate choices, documented in the eval README:

- `services:` is omitted. With it set, the sandbox shim appends
  `-x gotrue,kong,...` to `supabase start`, and the managed backend only
  accepts capability names (rest, auth, ...), so it rejects the command.
- Nothing tells the agent about the managed stack or its feature flag.
  Worktree isolation only exists on the managed stack (beta channel,
  gated behind experimental.stack); discovering that is part of what is
  measured. Expect every arm to fail until agents find it or the CLI
  flips the default.
- Environment-agnostic, like build-database-002-stack-lifecycle: the
  scorer never branches on which experiment arm it runs under.

Pure helpers are unit-tested in scoring.test.ts (run hint at the top of
the file); the shapes in the fixtures were captured from real beta CLI
output.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ss arms

The -cli-nodaemon and -cli-absent experiments restrict themselves to an
allowlist of eval ids so the forced-broken-Docker sandbox does not run
every CLI eval. Add the worktree-stacks eval to both: CLI-2400 asks for a
Docker-available and a Docker-less arm, and native-mode fleet startup
time is part of the launch story this eval reports on.

The pinned codex-gpt-5.6-luna experiment and the -cli-stable / -cli-beta
arms already pick the eval up through their suite / interface rules, so
no other experiment changes are needed. No skipEval is added anywhere:
the pinned CLI can only pass through per-worktree port and project_id
surgery, which is real signal about agent behavior without the feature.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@kanadgupta
kanadgupta requested a review from a team as a code owner September 15, 2026 20:30
@kanadgupta kanadgupta added the run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes label Sep 15, 2026
@vercel

vercel Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
evals Ready Ready Preview Sep 15, 2026 8:37pm UTC

Request Review

@kanadgupta

Copy link
Copy Markdown
Member Author

Superseded by #307, which rebuilds this on top of #281's rebased cli suite and needsDocker field, and fixes the two harness defects the CI run here exposed (prompt made 12/18 agents ask for a repo; scorer missed worktrees created outside the workspace).

@kanadgupta kanadgupta closed this Sep 17, 2026
@kanadgupta
kanadgupta deleted the kanad-claude/cli-2400-worktree-stacks branch September 17, 2026 20:49
Coly010 pushed a commit that referenced this pull request Sep 18, 2026
…2400)

One human plus N coding agents, each in its own git worktree, is the
headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree
needs its own isolated local Supabase stack with zero leakage between
them. Nothing in this repo measured that.

Scenario: starting from an empty sandbox, the agent creates a repo with
worktrees feature-a/b/c, starts a local stack in each, and adds a
different table (widgets/gadgets/gizmos) plus one seed row per worktree.
No seed data and no framework changes: the agent builds everything,
including the worktrees.

Scorer (end state only, all via ctx.exec because the harness's built-in
ctx.query/stackStatus resolve a single stack from the workspace root):

- three real git worktrees on distinct branches, discovered by finding a
  repo in the workspace and asking `git worktree list --porcelain`, so
  worktrees the agent placed outside the workspace still count and three
  plain directories don't;
- one live stack per worktree with three distinct host:port database
  endpoints, which is what catches aliasing/reuse;
- per-table schema isolation via to_regclass against all three stacks;
- at least one row per table in its home stack;
- each table created by a migration file in its worktree;
- an always-passing metrics check reporting fleet wall-clock across the
  three stacks, CLI version/channel, per-stack backend and runtime, and
  the number of `supabase start` invocations.

Stack resolution asks the managed backend first
(SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back
to the legacy one, the same shape as build-database-002-stack-lifecycle.
Verified against CLI 2.118.0-beta.37: the managed backend rejects the
legacy -o flag, the DB password is random per stack, and the legacy
backend cannot see managed stacks.

Suite placement follows #281: `evals/cli/` with `needsDocker: false`, so
the eval runs on the pinned Codex Luna baseline plus the four
`codex-gpt-5.6-luna-cli-*` arms with no experiment edits. `services:` is
omitted because the sandbox shim's legacy `-x` names are rejected by the
managed backend.

Two details come from the first CI run of an earlier draft (#295, 18
runs across six experiments):

- The prompt now says the project is brand-new and asks for the worktrees
  as folders in this directory. The earlier "set this repo up" made 12 of
  18 agents look for a repository in the empty workspace and stop to ask
  for one within about ten seconds.
- Worktree discovery goes through git rather than a name search under the
  workspace. Two agents built everything correctly with worktrees at
  /tmp/feature-*, and the name search reported them as missing.

The two runs that passed did so on the legacy backend by hand-editing
project_id and every port per worktree; the metrics check records the
backend so that path stays distinguishable from a managed-stack pass.

Pure helpers are unit-tested in scoring.test.ts (run hint at the top of
the file); the shapes in the fixtures were captured from real beta CLI
output. The full scorer was also run against three hand-built native
stacks on the beta CLI, with worktrees both inside and outside the
workspace, and passes all checks.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Coly010 pushed a commit that referenced this pull request Sep 23, 2026
…2400)

One human plus N coding agents, each in its own git worktree, is the
headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree
needs its own isolated local Supabase stack with zero leakage between
them. Nothing in this repo measured that.

Scenario: starting from an empty sandbox, the agent creates a repo with
worktrees feature-a/b/c, starts a local stack in each, and adds a
different table (widgets/gadgets/gizmos) plus one seed row per worktree.
No seed data and no framework changes: the agent builds everything,
including the worktrees.

Scorer (end state only, all via ctx.exec because the harness's built-in
ctx.query/stackStatus resolve a single stack from the workspace root):

- three real git worktrees on distinct branches, discovered by finding a
  repo in the workspace and asking `git worktree list --porcelain`, so
  worktrees the agent placed outside the workspace still count and three
  plain directories don't;
- one live stack per worktree with three distinct host:port database
  endpoints, which is what catches aliasing/reuse;
- per-table schema isolation via to_regclass against all three stacks;
- at least one row per table in its home stack;
- each table created by a migration file in its worktree;
- an always-passing metrics check reporting fleet wall-clock across the
  three stacks, CLI version/channel, per-stack backend and runtime, and
  the number of `supabase start` invocations.

Stack resolution asks the managed backend first
(SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back
to the legacy one, the same shape as build-database-002-stack-lifecycle.
Verified against CLI 2.118.0-beta.37: the managed backend rejects the
legacy -o flag, the DB password is random per stack, and the legacy
backend cannot see managed stacks.

Suite placement follows #281: `evals/cli/` with `needsDocker: false`, so
the eval runs on the pinned Codex Luna baseline plus the four
`codex-gpt-5.6-luna-cli-*` arms with no experiment edits. `services:` is
omitted because the sandbox shim's legacy `-x` names are rejected by the
managed backend.

Two details come from the first CI run of an earlier draft (#295, 18
runs across six experiments):

- The prompt now says the project is brand-new and asks for the worktrees
  as folders in this directory. The earlier "set this repo up" made 12 of
  18 agents look for a repository in the empty workspace and stop to ask
  for one within about ten seconds.
- Worktree discovery goes through git rather than a name search under the
  workspace. Two agents built everything correctly with worktrees at
  /tmp/feature-*, and the name search reported them as missing.

The two runs that passed did so on the legacy backend by hand-editing
project_id and every port per worktree; the metrics check records the
backend so that path stays distinguishable from a managed-stack pass.

Pure helpers are unit-tested in scoring.test.ts (run hint at the top of
the file); the shapes in the fixtures were captured from real beta CLI
output. The full scorer was also run against three hand-built native
stacks on the beta CLI, with worktrees both inside and outside the
workspace, and passes all checks.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 3e686664 Deployed Sep 15, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants