feat(evals): add build-cli-004-worktree-stacks (CLI-2400) - #295
Closed
kanadgupta wants to merge 7 commits into
Closed
kanadgupta wants to merge 7 commits into
kanadgupta wants to merge 7 commits into
Conversation
…periments Three regression evals for CLI-2398 share one prompt (init a project, start the local stack, add a seeded `notes` table) and one scorer, differing only in the sandbox's Docker state declared beside PROMPT.md in sandbox-environment.json: 002 is the Docker-available control, 003 has the docker client but no usable daemon, 004 has no docker binary at all. The scorer checks stack readiness, the seeded rows, zero container-runtime detours (install/start/escalation attempts), records cliVersion, resolvedRuntime, timeToReadyMs and raw-socket probes as metrics, and asks an LLM judge whether the agent named the real blocker when it failed. The two Docker-less arms fail today by design; they track the gap the CLI's native managed stack is meant to close. Two experiments clone codex-gpt-5.6-luna but install the latest stable or latest beta Supabase CLI (resolved lazily from npm dist-tags) via an experiment-land LocalStackRuntime built from the sandbox package's exports, so the pinned 2.67.1 baseline, stable and beta can be compared nightly across every interface: cli regression eval. The Docker-less arms are staged with DOCKER_HOST pointed at an unbound port plus root-owned PATH shims (the CI sandbox makes the socket world-writable, so permissions alone cannot block it); for 003 the docker shim still answers --version so the managed stack's runtime probe selects Docker as it would on a real host. The existing Luna experiments skip evals that need a Docker-less sandbox. Unit tests for the pure helpers run with: pnpm --filter @supabase-evals/framework exec vitest run --root ../.. experiments/_lib evals/build-database-002-stack-lifecycle Refs: https://linear.app/supabase/issue/CLI-2398/add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle
…-2398) An eval told in advance which environment it runs in, and graded toward that environment's expected story, measures conformance to an answer key rather than adaptation — and the marker-file design forced skipEval patches onto two unrelated pinned experiments just to suppress meaningless nightly fails. Collapse to one eval, build-database-002-stack-lifecycle, whose scorer asserts only environment-agnostic criteria (stack ready however it got there, seeded rows verified, zero container-runtime detours, a truthful final report) and reports resolvedRuntime as an observed metric. The forced environments become experiment variants: -cli-nodaemon (docker client present, daemon unreachable) and -cli-absent (no docker at all) wrap dockerAwareLocalStackRuntime with a `docker` option instead of reading an eval-side sandbox-environment.json, on the beta channel where the managed stack's Docker-less path lives, scoped to this scenario. The existing Luna experiments return to their upstream content and run the eval under their stock sandbox like any other. Also mirrors the upstream promptAddendum/buildSkillsPrompt API change (cd5b8d0, 03467fb): CLI agents get an empty addendum via buildToolSurfaceAddendum, matching the stock runtime. Refs: https://linear.app/supabase/issue/CLI-2398/add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle
…folder main now discovers evals as `evals/<suite>/<id>/` (#282) and derives the suite from the folder name, parsing every top-level directory under `evals/` as a suite. With main merged into this branch, the flat `evals/build-database-002-stack-lifecycle/` folder from PR #281 would make `evalSuiteSchema.parse('build-database-002-stack-lifecycle')` throw and abort discovery for every run. - Move the eval to `evals/regression/`, matching the `suite: regression` it already declared. - Drop the `suite:` frontmatter line: the runner ignores it in favor of the folder, so keeping it would be a second source of truth. - Update the vitest run hints and the schema cross-reference comment in `experiments/_lib` to the new path. The experiment allowlists reference the eval by id only, so they are unaffected. This commit is expected to become redundant once PR #281 rebases onto main. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One human plus N coding agents, each in its own git worktree, is the headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree needs its own isolated local Supabase stack with zero leakage between them. Nothing in this repo measured that. Scenario: the agent sets up a repo with worktrees feature-a/b/c, starts a local stack in each, and adds a different table (widgets/gadgets/gizmos) plus one seed row per worktree. No seed data: the agent builds everything, including the worktrees, so `local/` is not needed and no framework setup hook is required. Scorer (end state only, all via ctx.exec because the harness's built-in ctx.query/stackStatus resolve a single stack from the workspace root): - three real git worktrees on distinct branches (`git worktree list --porcelain`), so three plain directories don't count; - one live stack per worktree with three distinct host:port database endpoints, which is what catches aliasing/reuse; - per-table schema isolation via to_regclass against all three stacks; - at least one row per table in its home stack; - each table created by a migration file in its worktree; - an always-passing metrics check reporting fleet wall-clock across the three stacks, CLI version/channel, per-stack backend and runtime, and the number of `supabase start` invocations. Stack resolution asks the managed backend first (SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back to the legacy one (SUPABASE_EXPERIMENTAL_STACK=0 supabase status -o json). Verified against CLI 2.118.0-beta.37: the managed backend rejects the legacy -o flag outright, the DB password is random per stack, and the legacy backend cannot see managed stacks, so neither probe alone works. Deliberate choices, documented in the eval README: - `services:` is omitted. With it set, the sandbox shim appends `-x gotrue,kong,...` to `supabase start`, and the managed backend only accepts capability names (rest, auth, ...), so it rejects the command. - Nothing tells the agent about the managed stack or its feature flag. Worktree isolation only exists on the managed stack (beta channel, gated behind experimental.stack); discovering that is part of what is measured. Expect every arm to fail until agents find it or the CLI flips the default. - Environment-agnostic, like build-database-002-stack-lifecycle: the scorer never branches on which experiment arm it runs under. Pure helpers are unit-tested in scoring.test.ts (run hint at the top of the file); the shapes in the fixtures were captured from real beta CLI output. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ss arms The -cli-nodaemon and -cli-absent experiments restrict themselves to an allowlist of eval ids so the forced-broken-Docker sandbox does not run every CLI eval. Add the worktree-stacks eval to both: CLI-2400 asks for a Docker-available and a Docker-less arm, and native-mode fleet startup time is part of the launch story this eval reports on. The pinned codex-gpt-5.6-luna experiment and the -cli-stable / -cli-beta arms already pick the eval up through their suite / interface rules, so no other experiment changes are needed. No skipEval is added anywhere: the pinned CLI can only pass through per-worktree port and project_id surgery, which is real signal about agent behavior without the feature. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
6 tasks
5 tasks
Member
Author
Coly010
pushed a commit
that referenced
this pull request
Sep 18, 2026
…2400) One human plus N coding agents, each in its own git worktree, is the headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree needs its own isolated local Supabase stack with zero leakage between them. Nothing in this repo measured that. Scenario: starting from an empty sandbox, the agent creates a repo with worktrees feature-a/b/c, starts a local stack in each, and adds a different table (widgets/gadgets/gizmos) plus one seed row per worktree. No seed data and no framework changes: the agent builds everything, including the worktrees. Scorer (end state only, all via ctx.exec because the harness's built-in ctx.query/stackStatus resolve a single stack from the workspace root): - three real git worktrees on distinct branches, discovered by finding a repo in the workspace and asking `git worktree list --porcelain`, so worktrees the agent placed outside the workspace still count and three plain directories don't; - one live stack per worktree with three distinct host:port database endpoints, which is what catches aliasing/reuse; - per-table schema isolation via to_regclass against all three stacks; - at least one row per table in its home stack; - each table created by a migration file in its worktree; - an always-passing metrics check reporting fleet wall-clock across the three stacks, CLI version/channel, per-stack backend and runtime, and the number of `supabase start` invocations. Stack resolution asks the managed backend first (SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back to the legacy one, the same shape as build-database-002-stack-lifecycle. Verified against CLI 2.118.0-beta.37: the managed backend rejects the legacy -o flag, the DB password is random per stack, and the legacy backend cannot see managed stacks. Suite placement follows #281: `evals/cli/` with `needsDocker: false`, so the eval runs on the pinned Codex Luna baseline plus the four `codex-gpt-5.6-luna-cli-*` arms with no experiment edits. `services:` is omitted because the sandbox shim's legacy `-x` names are rejected by the managed backend. Two details come from the first CI run of an earlier draft (#295, 18 runs across six experiments): - The prompt now says the project is brand-new and asks for the worktrees as folders in this directory. The earlier "set this repo up" made 12 of 18 agents look for a repository in the empty workspace and stop to ask for one within about ten seconds. - Worktree discovery goes through git rather than a name search under the workspace. Two agents built everything correctly with worktrees at /tmp/feature-*, and the name search reported them as missing. The two runs that passed did so on the legacy backend by hand-editing project_id and every port per worktree; the metrics check records the backend so that path stays distinguishable from a managed-stack pass. Pure helpers are unit-tested in scoring.test.ts (run hint at the top of the file); the shapes in the fixtures were captured from real beta CLI output. The full scorer was also run against three hand-built native stacks on the beta CLI, with worktrees both inside and outside the workspace, and passes all checks. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Coly010
pushed a commit
that referenced
this pull request
Sep 23, 2026
…2400) One human plus N coding agents, each in its own git worktree, is the headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree needs its own isolated local Supabase stack with zero leakage between them. Nothing in this repo measured that. Scenario: starting from an empty sandbox, the agent creates a repo with worktrees feature-a/b/c, starts a local stack in each, and adds a different table (widgets/gadgets/gizmos) plus one seed row per worktree. No seed data and no framework changes: the agent builds everything, including the worktrees. Scorer (end state only, all via ctx.exec because the harness's built-in ctx.query/stackStatus resolve a single stack from the workspace root): - three real git worktrees on distinct branches, discovered by finding a repo in the workspace and asking `git worktree list --porcelain`, so worktrees the agent placed outside the workspace still count and three plain directories don't; - one live stack per worktree with three distinct host:port database endpoints, which is what catches aliasing/reuse; - per-table schema isolation via to_regclass against all three stacks; - at least one row per table in its home stack; - each table created by a migration file in its worktree; - an always-passing metrics check reporting fleet wall-clock across the three stacks, CLI version/channel, per-stack backend and runtime, and the number of `supabase start` invocations. Stack resolution asks the managed backend first (SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back to the legacy one, the same shape as build-database-002-stack-lifecycle. Verified against CLI 2.118.0-beta.37: the managed backend rejects the legacy -o flag, the DB password is random per stack, and the legacy backend cannot see managed stacks. Suite placement follows #281: `evals/cli/` with `needsDocker: false`, so the eval runs on the pinned Codex Luna baseline plus the four `codex-gpt-5.6-luna-cli-*` arms with no experiment edits. `services:` is omitted because the sandbox shim's legacy `-x` names are rejected by the managed backend. Two details come from the first CI run of an earlier draft (#295, 18 runs across six experiments): - The prompt now says the project is brand-new and asks for the worktrees as folders in this directory. The earlier "set this repo up" made 12 of 18 agents look for a repository in the empty workspace and stop to ask for one within about ten seconds. - Worktree discovery goes through git rather than a name search under the workspace. Two agents built everything correctly with worktrees at /tmp/feature-*, and the name search reported them as missing. The two runs that passed did so on the legacy backend by hand-editing project_id and every port per worktree; the metrics check records the backend so that path stays distinguishable from a managed-stack pass. Pure helpers are unit-tested in scoring.test.ts (run hint at the top of the file); the shapes in the fixtures were captured from real beta CLI output. The full scorer was also run against three hand-built native stacks on the beta CLI, with worktrees both inside and outside the workspace, and passes all checks. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tracks CLI-2400 under the Slim CLI evals RFC.
Note
Stacked on #281. This branch is Colum's branch plus
origin/mainmerged in plus the commits below. The diff will shrink to the last two commits once #281 lands.The problem it's solving
The headline agent workflow for the Slim CLI launch is one human plus N coding agents, each working in its own git worktree, each needing its own isolated local Supabase stack. Nothing in this repo measured whether an agent can actually get there: three worktrees, three stacks, zero leakage.
What the PR adds
One new regression eval,
evals/regression/build-cli-004-worktree-stacks/. The agent is asked to set a repo up with worktreesfeature-a,feature-b,feature-c, start a local stack in each, and add a different table (widgets/gadgets/gizmos) plus one sample row per worktree. No seed data: the agent builds everything, including the worktrees.A scorer that grades the end state, never the method. Because the harness's built-in
ctx.query/stackStatusresolve a single stack from the workspace root, every check addresses stacks itself bycd-ing into each worktree:to_regclassagainst all three stacks;metricscheck reporting fleet wall-clock across the three stacks, CLI version and channel, per-stack backend and runtime, and the number ofsupabase startinvocations.The scorer asks the managed backend first (
SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env --output-format json) and falls back to the legacy one, so it works whichever path the agent took and however it enabled the flag. Same environment-agnostic shape as #281: it never branches on which experiment arm it runs under.Two experiment edits. The eval id is added to the Docker-less allowlists in
-cli-nodaemonand-cli-absent. The pinned Codex Luna experiment and the-cli-stable/-cli-betaarms already pick it up through their suite and interface rules, so the eval runs under all five. NoskipEvalanywhere.One housekeeping commit.
mainnow discovers evals asevals/<suite>/<id>/and parses every top-level folder as a suite name, so #281's flatevals/build-database-002-stack-lifecycle/breaks discovery once main is merged in. This branch moves it intoevals/regression/and drops the now-ignoredsuite:frontmatter line. That commit becomes redundant when #281 rebases.Deliberate choices
experimental.stack). Whether agents can discover that is part of what's measured. Expect every arm to fail until they do or the CLI flips the default.services:is omitted. With it set, the sandbox shim appends legacy container names (-x gotrue,kong,...) tosupabase start, and the managed backend rejects those outright (Unknown stack capabilities in --exclude). This also affects feat(evals): add build-database-002-stack-lifecycle CLI eval (CLI-2398) #281's eval, which setsservices: [].supabase initstill pins ports inconfig.toml. The managed stack treats them as exact intents, so the second worktree's start fails withPersisted database port is unavailableuntil the ports are removed and become automatic. That is exactly the "no port surgery" promise the launch makes.Verification
Verified by hand against CLI 2.118.0-beta.37 on the native runtime: three sibling worktrees get three distinct stack identities with no config edits, and a table created in one is absent from the others. The managed backend rejects
supabase status -o json, the DB password is random per stack, and excludingauthon this beta fails start withPersisted API gateway material is incomplete.scoring.test.ts, 23 cases; fixtures captured from real beta CLI output)pnpm typecheck, Biomerun-evals-changedlabel on this PR.🤖 Generated with Claude Code