Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Coly010
added a commit
that referenced
this pull request
Sep 23, 2026
Final chunk of the closed #308 — the piece @mattrossman asked for: the minimal CLI suite plumbing and CODEOWNERS, standalone, with no CLI experiment internals to review. **15 files, 260 lines**, and four of those files are 11–13 lines each. Everything it depended on is now on `main` (#315's `experiments/<owner>/` layout, and #324 + #325's sandbox options), so this is based on `main` and the diff is only its own content. ## What it adds `cli` as an eval suite and an experiment suite, `@supabase/cli` ownership over `/evals/cli/`, `/experiments/cli/` and `cli-eval-results.json`, and the workflow wiring so `eval-refresh` and `append-gh-pages-history` know the suite exists. The substance is five environment columns: one scenario run unchanged across forced CLI environments, so a failure isolates to *which* environment broke. They're thin now that `experiments/presets.ts` exists — the whole of the Docker-less column is: ```ts export default defineExperiment({ ...codexGpt6Luna, suite: ['cli'], // beta: the Docker-less path only exists in the managed stack, which ships in beta. localStack: localStackRuntime({ cliVersion: 'beta', docker: 'absent' }), skipEval: skipUnlessDockerless, }); ``` | column | environment | picks up | |---|---|---| | `codex-gpt-6-luna-cli-pinned` | repo-pinned CLI, Docker available | every `interface: cli` eval that isn't hosted-linked | | `…-cli-stable` | npm `latest` | same | | `…-cli-beta` | npm `beta` | same | | `…-cli-nodaemon` | beta, daemon unreachable | also needs `needsDocker: false` + `projectRunning: false` | | `…-cli-absent` | beta, no `docker` binary | also needs `needsDocker: false` + `projectRunning: false` | All five columns share one `skipEval: skipUnlessCli`, so they run the same eval set and stay comparable — which is the whole point of the suite. An earlier revision made the pinned column the shared benchmark experiment with `'cli'` appended to its `suite`; that experiment has no `skipEval`, so it picked up hosted-linked evals the other four skip. A CLI-owned pinned experiment fixes that and leaves this PR touching nothing outside CLI-owned paths. `skipUnlessCli` / `skipUnlessDockerless` are the only things left over from the old `experiments/_lib/`; they live in `experiments/cli/lib/`, which discovery ignores for free since it only matches `*.experiment.ts` directly under an owner directory. ## Verified, including the parts tests can't reach `format:check`, `typecheck`, and the sandbox, core, framework, `test:cli-lib`, vercel-runner and web suites all pass. Two things unit tests can't cover, checked directly instead: - **Discovery**, since it's filename-convention-based and fails silently: all 16 experiments resolve and load, `experiments/cli/lib/` is correctly not treated as an experiment, and the five columns resolve to the intended runtimes and channels (`beta`, `beta`, `beta`, `stable`, and none for the pinned baseline). - **The workflow**, which only ever runs in CI, no longer resolves channels at all — it passes through the manual `cli_stable_version` / `cli_beta_version` dispatch inputs and otherwise lets `run-vercel-evals.ts` derive the channels from the actual pairs. Pre-filling both env vars would have short-circuited that narrowing. Each run's results now record the CLI version the sandbox actually installed, read from the session's environment marker, so a `stable` row names its binary rather than just its channel. That needed `cliVersionSchema` widened — it rejected every prerelease, so a beta run's real version could not have been recorded at all. `test:cli-lib` carries `--passWithNoTests` deliberately: `evals/cli/` has no scorer tests until #281/#314/#316 land, and without the flag the script's exit code depends on which of its two paths happens to be populated. ## Note on the column names These follow the `gpt-6-luna` rename from #328. They key `cli-eval-results.json`, and no results exist yet, so there is nothing to churn — but that also means the names should settle before the suite is first run. ## Next #281, #314, #316 and #307 get re-parented onto `main` once this lands — they currently point at the retained branch of the closed #308. /cc @Rodriguespn @kanadgupta
…399) Two independent local Supabase projects, client-a and client-b, running concurrently in one sandbox. The agent creates both, gives each a clients table naming that client, and reports which API port each landed on. Port collision between concurrent local projects is the scenario parallel instances exists to fix, and nothing measured whether an agent can actually drive it. The scorer asserts CLI-2399's three pass criteria plus zero detours: both project directories initialised (discovered, not assumed), both stacks reach ready, the two stacks resolve to distinct ports, each database holds only its own marker row, both real API ports appear in the report, and cliDetours is zero. Port attribution between the two projects is left to the judge, which grades a truthful success and a truthful failure alike. resolvedRuntime, timeToReadyMs and the observed ports are reported as metrics, never asserted, and the runtime marker type omits the experiment's docker field so the eval cannot grade against its own arm. The marker-isolation check runs both directions: client-a's database must contain client-a and must NOT contain client-b. A single stack holding both rows is the "addressed the wrong stack" failure this eval exists to catch, so a one-directional check would pass it. needsDocker: false and projectRunning: false, so it runs under all five cli arms. The harness starts no stack at all when projectRunning is false, so running two is entirely the agent's own work.
…arker() Replaces the local readRuntimeMarker/RuntimeMarker/RUNTIME_MARKER_PATH, which read the marker via a shell cat that an agent-writable PATH could shadow, with the framework's root-read, shape-validated ctx.environmentMarker(). Also trims stale provenance comments and narration to match the repo's comment policy.
Coly010
force-pushed
the
columferry/cli-2399-build-database-003-parallel-projects
branch
from
September 23, 2026 15:05
4c048d6 to
76c1ab0
Compare
Coly010
changed the base branch from
columferry/cli-2398-cli-suite-and-arms
to
main
September 23, 2026 15:05
5 tasks done
…uild-database-003-parallel-projects
… into columferry/cli-2399-build-database-003-parallel-projects
…o evals/cli/lib Split the single scoring module into EVAL.ts plus projects/stacks/report/ metrics modules on the shared evals/cli/lib helpers, and add a README. Behaviour fixes: - reported api ports: match each project's API port as a whole number, not inside a longer one - reported api ports: fail with "stack resolved via <backend> but reported no API URL" instead of a generic unavailable note - marker isolation is case-insensitive in both directions - project discovery prefers an exact basename (client-a over client-a-old) and fails when both names resolve to the same directory - container-runtime detours are judged by the shared detour judge instead of the regex count, which is now a metric only - PROMPT.md drops the unused migrations topic and cites discussion #5968
… into columferry/cli-2399-build-database-003-parallel-projects
… into columferry/cli-2399-build-database-003-parallel-projects
6 tasks done
… into columferry/cli-2399-build-database-003-parallel-projects
Coly010
added a commit
that referenced
this pull request
Oct 1, 2026
…354) ## Current Behavior `build-database-002-stack-lifecycle` owns its detour parser, detour judge rubric, stack resolution and metrics readers. The next two CLI evals (#314 `build-database-003-parallel-projects`, #316 `resolve-database-002-stale-stack-cleanup`) each carry a byte-identical copy of that code (~700 lines apiece), so every detour-policy or probe fix would need landing three times. ## Expected Behavior Adds `evals/cli/lib/`, colocated helpers for CLI-team evals, following the same convention as `experiments/cli/lib/` discussed in #team-ai. Eval discovery only picks up folders with a `PROMPT.md`, so `lib/` is never treated as an eval. `lib/README.md` states the rule: lib holds probe helpers, shared judge policy text and input formatters; each eval's `EVAL.ts` still owns check composition, every `judge()` call, its scenario rubric and the default export. | Module | What it holds | |---|---| | `detours.ts` | the regex/shell-parsing diagnostics (moved from 002 unchanged), `formatDetourJudgeInput`, `DETOUR_CHECK_NAME`, `detourJudgeRubric(scenario)` | | `stack.ts` | one `resolveStack(ctx, target)` for a repo-root project (002), a per-directory project (#314) and a named managed stack (#316): managed-named → managed → legacy, first `DB_URL` wins; plus JSON/URL readers and the `select 1` readiness probe | | `metrics.ts` | CLI version, session start, `DOCKER_HOST` clears | | `report.ts` | `formatGroundTruthJudgeInput` for truthful-report judges, `mentionsNumber` (digit-boundary port matching) | | `projects.ts` | per-name project discovery via `*/supabase/config.toml` (exact basename preferred) | | `markers.ts` | marker-row reader and N-way `checkMarkerIsolation` (each db holds its own label and none of the others) | | `cli-invocations.ts` | argv-parsed `supabase` invocations the agent actually ran (skips echoed/passive text), with `cd`/`--workdir`/`--stack`/`--project-id` target attribution and shell-loop (`$var`) detection | `projects.ts`, `markers.ts` and `cli-invocations.ts` have no consumer on `main` yet; #314 and #316 use them and are based on this branch. **No behaviour change for `build-database-002`:** check names, probe command strings, failure notes, metrics keys and judge inputs are the same. Tests pin this: the detour rubric matches 002's previous rubric (modulo line wrapping), the ground-truth judge input is byte-identical, and `resolveStack`'s root target issues exactly 002's commands with no `cd`. SQL comment/literal masking stays in 002, since nothing else matches SQL text. Lib files are typechecked through 002's `EVAL.ts` import graph; `apps/framework/tsconfig.json` is untouched. ## Test plan - [x] `vitest run --root ../.. evals/cli` (lib + 002) - [x] `apps/framework` `tsc --noEmit` - [x] `pnpm eval:dry -- --suite cli --experiment-suite cli` → 5 experiments × 1 eval, `lib` not listed - [x] `pnpm format:check` - [x] `pnpm check`, all steps except the judge smoke step, which needs a provider API key locally 🤖 Generated with [Claude Code](https://claude.com/claude-code)
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Current Behavior
Port collisions between concurrent local Supabase projects are a long-running pain point (discussion #5968). The CLI historically assumed one machine means one fixed port set. Running parallel stacks is the flagship behaviour that work is meant to fix, and nothing measures whether an agent can actually drive it.
Expected Behavior
Adds
evals/cli/build-database-003-parallel-projects(interface: cli,projectRunning: false,needsDocker: false). The agent stands up two independent local Supabase projects,client-aandclient-b, concurrently in one sandbox. Each project gets aclientstable holding a row naming that client, and the agent reports which API port each one landed on. The prompt never names a command, a port or Docker.Structured like
build-database-002:EVAL.tscomposes the checks and owns bothjudge()calls;projects.ts,stacks.ts,report.tsandmetrics.tshold the deterministic checks, each with its own tests. Shared probing (stack resolution, project discovery, marker reads, detour policy) comes fromevals/cli/lib. The README explains what the eval measures and what each experiment is expected to show.*/supabase/config.toml, exact basename preferred (client-a-olddoesn't makeclient-aambiguous)select 1client-a's db hasclient-aand notclient-b, and vice versacliDetours(regex diagnostic) — reported, never assertedKnown-bad cases covered by unit tests:
cd client-bresolving toclient-a's stackconfig.tomlclientstable with three rows, or noneThe CLI-2399 issue describes this eval as blocked on multi-stack session support. That no longer applies: with
projectRunning: falsethe harness starts zero stacks and only installs the CLI, so running two stacks is entirely the agent's job, which is the point of the scenario.Observed in CI so far:
absentfails 0/3 by design; the detour and truthful-report checks pass 3/3.nodaemonran under a harness PATH bug that made it behave likeabsent(fixed in fix(sandbox): resolve shims in login shells #355). Results will be refreshed after fix(sandbox): resolve shims in login shells #355/feat(core): record the working directory of agent tool calls #356 merge.API_URLhasn't been seen live yet, though the beta CLI exports it (apps/cli/src/commands/experimental/stack/status/status.env.ts).Related Issue(s)
Refs https://linear.app/supabase/issue/CLI-2399
Test plan
vitest run --root ../.. evals/cliapps/frameworktsc --noEmitpnpm eval:dry -- --suite cli --experiment-suite cli→ 003 planned under all 5 cli experimentspnpm format:checkpnpm check, all steps except the judge smoke step, which needs a provider API key locallyrun-evals-changed: inspected passing runs and each distinct failure shapenodaemon🤖 Generated with Claude Code