Skip to content

feat(evals): add build-database-003-parallel-projects CLI eval (CLI-2399) - #314

Open
Coly010 wants to merge 18 commits into
columferry/cli-evals-shared-libfrom
columferry/cli-2399-build-database-003-parallel-projects
Open

Coly010 wants to merge 18 commits into
columferry/cli-evals-shared-libfrom
columferry/cli-2399-build-database-003-parallel-projects

Conversation

@Coly010

@Coly010 Coly010 commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Current Behavior

Port collisions between concurrent local Supabase projects are a long-running pain point (discussion #5968). The CLI historically assumed one machine means one fixed port set. Running parallel stacks is the flagship behaviour that work is meant to fix, and nothing measures whether an agent can actually drive it.

Expected Behavior

Adds evals/cli/build-database-003-parallel-projects (interface: cli, projectRunning: false, needsDocker: false). The agent stands up two independent local Supabase projects, client-a and client-b, concurrently in one sandbox. Each project gets a clients table holding a row naming that client, and the agent reports which API port each one landed on. The prompt never names a command, a port or Docker.

Base is #354 (evals/cli/lib) so this diff shows only the eval; it retargets to main once #354 merges.

Structured like build-database-002: EVAL.ts composes the checks and owns both judge() calls; projects.ts, stacks.ts, report.ts and metrics.ts hold the deterministic checks, each with its own tests. Shared probing (stack resolution, project discovery, marker reads, detour policy) comes from evals/cli/lib. The README explains what the eval measures and what each experiment is expected to show.

Check Proves
two projects initialised both project dirs found via */supabase/config.toml, exact basename preferred (client-a-old doesn't make client-a ambiguous)
both stacks reach ready each project's stack resolves from its own directory and answers select 1
stacks are on distinct ports two independent stacks, not one reused or clobbered
each project holds only its own marker row both directions, case-insensitive: client-a's db has client-a and not client-b, and vice versa
each clients table holds exactly one row the prompt asks for a single row per client; duplicates or an empty table fail
reported api ports match the running stacks each project's real API port appears in the final report as a whole number
no container-runtime detours LLM judge over the complete executed commands only (shared rubric with 002)
metrics per-project backend/runtime, ports, time to ready, cliDetours (regex diagnostic) — reported, never asserted
final report is truthful about both projects LLM judge given harness ground truth plus the transcript with tool outputs; fails on a misstated outcome, rows or ports, or an omitted/swapped API port, not on rows the report didn't recite

Known-bad cases covered by unit tests:

  • a port that only appears inside a longer number
  • rows swapped between the two databases
  • one database holding both rows
  • cd client-b resolving to client-a's stack
  • a stack that resolves but reports no API URL, which fails with a clear note instead of falling back to config.toml
  • a clients table with three rows, or none

The CLI-2399 issue describes this eval as blocked on multi-stack session support. That no longer applies: with projectRunning: false the harness starts zero stacks and only installs the CLI, so running two stacks is entirely the agent's job, which is the point of the scenario.

Observed in CI so far:

⚠️ All passing runs resolved via the legacy backend. The managed backend's API_URL hasn't been seen live yet, though the beta CLI exports it (apps/cli/src/commands/experimental/stack/status/status.env.ts).

Related Issue(s)

Refs https://linear.app/supabase/issue/CLI-2399

Test plan

🤖 Generated with Claude Code

@Coly010
Coly010 requested a review from a team as a code owner September 21, 2026 10:00
@vercel

vercel Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
evals Ready Ready Preview Oct 1, 2026 3:04pm UTC

Request Review

Coly010 added a commit that referenced this pull request Sep 23, 2026
Final chunk of the closed #308 — the piece @mattrossman asked for: the
minimal CLI suite plumbing and CODEOWNERS, standalone, with no CLI
experiment internals to review. **15 files, 260 lines**, and four of
those files are 11–13 lines each.

Everything it depended on is now on `main` (#315's
`experiments/<owner>/` layout, and #324 + #325's sandbox options), so
this is based on `main` and the diff is only its own content.

## What it adds

`cli` as an eval suite and an experiment suite, `@supabase/cli`
ownership over `/evals/cli/`, `/experiments/cli/` and
`cli-eval-results.json`, and the workflow wiring so `eval-refresh` and
`append-gh-pages-history` know the suite exists.

The substance is five environment columns: one scenario run unchanged
across forced CLI environments, so a failure isolates to *which*
environment broke. They're thin now that `experiments/presets.ts` exists
— the whole of the Docker-less column is:

```ts
export default defineExperiment({
  ...codexGpt6Luna,
  suite: ['cli'],
  // beta: the Docker-less path only exists in the managed stack, which ships in beta.
  localStack: localStackRuntime({ cliVersion: 'beta', docker: 'absent' }),
  skipEval: skipUnlessDockerless,
});
```

| column | environment | picks up |
|---|---|---|
| `codex-gpt-6-luna-cli-pinned` | repo-pinned CLI, Docker available |
every `interface: cli` eval that isn't hosted-linked |
| `…-cli-stable` | npm `latest` | same |
| `…-cli-beta` | npm `beta` | same |
| `…-cli-nodaemon` | beta, daemon unreachable | also needs `needsDocker:
false` + `projectRunning: false` |
| `…-cli-absent` | beta, no `docker` binary | also needs `needsDocker:
false` + `projectRunning: false` |

All five columns share one `skipEval: skipUnlessCli`, so they run the
same eval set and stay comparable — which is the whole point of the
suite. An earlier revision made the pinned column the shared benchmark
experiment with `'cli'` appended to its `suite`; that experiment has no
`skipEval`, so it picked up hosted-linked evals the other four skip. A
CLI-owned pinned experiment fixes that and leaves this PR touching
nothing outside CLI-owned paths.

`skipUnlessCli` / `skipUnlessDockerless` are the only things left over
from the old `experiments/_lib/`; they live in `experiments/cli/lib/`,
which discovery ignores for free since it only matches `*.experiment.ts`
directly under an owner directory.

## Verified, including the parts tests can't reach

`format:check`, `typecheck`, and the sandbox, core, framework,
`test:cli-lib`, vercel-runner and web suites all pass.

Two things unit tests can't cover, checked directly instead:

- **Discovery**, since it's filename-convention-based and fails
silently: all 16 experiments resolve and load, `experiments/cli/lib/` is
correctly not treated as an experiment, and the five columns resolve to
the intended runtimes and channels (`beta`, `beta`, `beta`, `stable`,
and none for the pinned baseline).
- **The workflow**, which only ever runs in CI, no longer resolves
channels at all — it passes through the manual `cli_stable_version` /
`cli_beta_version` dispatch inputs and otherwise lets
`run-vercel-evals.ts` derive the channels from the actual pairs.
Pre-filling both env vars would have short-circuited that narrowing.

Each run's results now record the CLI version the sandbox actually
installed, read from the session's environment marker, so a `stable` row
names its binary rather than just its channel. That needed
`cliVersionSchema` widened — it rejected every prerelease, so a beta
run's real version could not have been recorded at all.

`test:cli-lib` carries `--passWithNoTests` deliberately: `evals/cli/`
has no scorer tests until #281/#314/#316 land, and without the flag the
script's exit code depends on which of its two paths happens to be
populated.

## Note on the column names

These follow the `gpt-6-luna` rename from #328. They key
`cli-eval-results.json`, and no results exist yet, so there is nothing
to churn — but that also means the names should settle before the suite
is first run.

## Next

#281, #314, #316 and #307 get re-parented onto `main` once this lands —
they currently point at the retained branch of the closed #308. /cc
@Rodriguespn @kanadgupta
…399)

Two independent local Supabase projects, client-a and client-b, running
concurrently in one sandbox. The agent creates both, gives each a clients
table naming that client, and reports which API port each landed on.

Port collision between concurrent local projects is the scenario parallel
instances exists to fix, and nothing measured whether an agent can
actually drive it.

The scorer asserts CLI-2399's three pass criteria plus zero detours: both
project directories initialised (discovered, not assumed), both stacks
reach ready, the two stacks resolve to distinct ports, each database
holds only its own marker row, both real API ports appear in the report,
and cliDetours is zero. Port attribution between the two projects is left
to the judge, which grades a truthful success and a truthful failure
alike. resolvedRuntime, timeToReadyMs and the observed ports are reported
as metrics, never asserted, and the runtime marker type omits the
experiment's docker field so the eval cannot grade against its own arm.

The marker-isolation check runs both directions: client-a's database must
contain client-a and must NOT contain client-b. A single stack holding
both rows is the "addressed the wrong stack" failure this eval exists to
catch, so a one-directional check would pass it.

needsDocker: false and projectRunning: false, so it runs under all five
cli arms. The harness starts no stack at all when projectRunning is
false, so running two is entirely the agent's own work.
…arker()

Replaces the local readRuntimeMarker/RuntimeMarker/RUNTIME_MARKER_PATH,
which read the marker via a shell cat that an agent-writable PATH could
shadow, with the framework's root-read, shape-validated
ctx.environmentMarker(). Also trims stale provenance comments and
narration to match the repo's comment policy.
@Coly010
Coly010 force-pushed the columferry/cli-2399-build-database-003-parallel-projects branch from 4c048d6 to 76c1ab0 Compare September 23, 2026 15:05
@Coly010
Coly010 changed the base branch from columferry/cli-2398-cli-suite-and-arms to main September 23, 2026 15:05
… into columferry/cli-2399-build-database-003-parallel-projects
…o evals/cli/lib

Split the single scoring module into EVAL.ts plus projects/stacks/report/
metrics modules on the shared evals/cli/lib helpers, and add a README.

Behaviour fixes:
- reported api ports: match each project's API port as a whole number, not
  inside a longer one
- reported api ports: fail with "stack resolved via <backend> but reported
  no API URL" instead of a generic unavailable note
- marker isolation is case-insensitive in both directions
- project discovery prefers an exact basename (client-a over client-a-old)
  and fails when both names resolve to the same directory
- container-runtime detours are judged by the shared detour judge instead
  of the regex count, which is now a metric only
- PROMPT.md drops the unused migrations topic and cites discussion #5968
@Coly010
Coly010 changed the base branch from main to columferry/cli-evals-shared-lib October 1, 2026 09:56
@Coly010 Coly010 added the run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes label Oct 1, 2026
Coly010 and others added 2 commits October 1, 2026 15:54
Coly010 added a commit that referenced this pull request Oct 1, 2026
…354)

## Current Behavior

`build-database-002-stack-lifecycle` owns its detour parser, detour
judge rubric, stack resolution and metrics readers. The next two CLI
evals (#314 `build-database-003-parallel-projects`, #316
`resolve-database-002-stale-stack-cleanup`) each carry a byte-identical
copy of that code (~700 lines apiece), so every detour-policy or probe
fix would need landing three times.

## Expected Behavior

Adds `evals/cli/lib/`, colocated helpers for CLI-team evals, following
the same convention as `experiments/cli/lib/` discussed in #team-ai.
Eval discovery only picks up folders with a `PROMPT.md`, so `lib/` is
never treated as an eval. `lib/README.md` states the rule: lib holds
probe helpers, shared judge policy text and input formatters; each
eval's `EVAL.ts` still owns check composition, every `judge()` call, its
scenario rubric and the default export.

| Module | What it holds |
|---|---|
| `detours.ts` | the regex/shell-parsing diagnostics (moved from 002
unchanged), `formatDetourJudgeInput`, `DETOUR_CHECK_NAME`,
`detourJudgeRubric(scenario)` |
| `stack.ts` | one `resolveStack(ctx, target)` for a repo-root project
(002), a per-directory project (#314) and a named managed stack (#316):
managed-named → managed → legacy, first `DB_URL` wins; plus JSON/URL
readers and the `select 1` readiness probe |
| `metrics.ts` | CLI version, session start, `DOCKER_HOST` clears |
| `report.ts` | `formatGroundTruthJudgeInput` for truthful-report
judges, `mentionsNumber` (digit-boundary port matching) |
| `projects.ts` | per-name project discovery via
`*/supabase/config.toml` (exact basename preferred) |
| `markers.ts` | marker-row reader and N-way `checkMarkerIsolation`
(each db holds its own label and none of the others) |
| `cli-invocations.ts` | argv-parsed `supabase` invocations the agent
actually ran (skips echoed/passive text), with
`cd`/`--workdir`/`--stack`/`--project-id` target attribution and
shell-loop (`$var`) detection |

`projects.ts`, `markers.ts` and `cli-invocations.ts` have no consumer on
`main` yet; #314 and #316 use them and are based on this branch.

**No behaviour change for `build-database-002`:** check names, probe
command strings, failure notes, metrics keys and judge inputs are the
same. Tests pin this: the detour rubric matches 002's previous rubric
(modulo line wrapping), the ground-truth judge input is byte-identical,
and `resolveStack`'s root target issues exactly 002's commands with no
`cd`. SQL comment/literal masking stays in 002, since nothing else
matches SQL text.

Lib files are typechecked through 002's `EVAL.ts` import graph;
`apps/framework/tsconfig.json` is untouched.

## Test plan

- [x] `vitest run --root ../.. evals/cli` (lib + 002)
- [x] `apps/framework` `tsc --noEmit`
- [x] `pnpm eval:dry -- --suite cli --experiment-suite cli` → 5
experiments × 1 eval, `lib` not listed
- [x] `pnpm format:check`
- [x] `pnpm check`, all steps except the judge smoke step, which needs a
provider API key locally

🤖 Generated with [Claude Code](https://claude.com/claude-code)

This branch was successfully deployed

1 active deployment
Preview — 9d1160cf Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant