Skip to content

Repository files navigation

Firstpass — cheap until proven otherwise. It routes every request to the cheapest model, gates the real output, and escalates only on proof of need. Measured on 974 real MBPP coding tasks: it recovered 15 points of quality over the cheap model alone and made zero of those 974 tasks worse, at 12–57% lower cost per success depending on how expensive your top rung is — with a distribution-free guarantee of ≤10% wrong answers served at 95% confidence and a signed receipt every call.

⚡ Firstpass

The verification layer for LLM serving. Nothing ships until a real check passes.

On 974 real coding tasks it recovered +15.2 points of quality over the cheap model alone — and made zero of those 974 tasks worse. At 12–57% lower cost per success, depending on how expensive your top rung is.

Your check — your tests, your schema, your judge — runs on every answer before it is served. What passes ships; what fails escalates and is checked again. Every decision leaves a hash-chained receipt you can re-derive, and wrong answers are capped by a distribution-free guarantee rather than a promise.

Routing falls out of that. Because the cheapest model clears the gate most of the time, you stop paying frontier prices for work that never needed them — but the cost saving is the consequence, not the mechanism.

CI release PyPI license stars

Website · Install · Quickstart · How it works · The proof · Docs


Proof over prediction

Almost every way of choosing a model decides before the output exists — a classifier reads your prompt and guesses, or a gateway picks on price and availability. Guess wrong and you find out in production, with no artifact explaining why.

Firstpass is not in that race. It decides after the output exists, which is the only point at which the question "is this good enough to serve?" can actually be answered.

Firstpass decides by proof. It opens on the cheapest model in your ladder, then gates the actual output — runs your tests, checks a schema, asks a judge, or measures self-consistency. Pass → it serves. Fail → it escalates exactly one rung and gates again. The cheap model handles most traffic; the frontier model is spent only when the cheap one is provably not enough. Every decision is a tamper-evident, hash-chained receipt you (or an auditor) can re-derive independently — and a distribution-free bound caps how often a wrong answer is served.

Cheap until proven otherwise. You pay frontier prices only when a real check proves you must.

The cascade-with-verification idea is not ours — FrugalGPT published it in 2023 and AutoMix refined it in 2024. What no paper or product ships is the receipt, the served-failure bound, and adding a model without retraining anything. What we borrow and what is new →

Where this fits, stated plainly. Verification needs something to verify against, so Firstpass is strongest where "correct" is checkable — code with tests, structured output against a schema, extraction, anything with an oracle. For open-ended prose the available gates are a judge or k-sample self-consistency, and on our own measurements a judge gate did not pay for itself once its calls were metered. If your output has no check, this is the wrong tool and we would rather say so here than after you have deployed it.

The measured consequence, on 974 real coding tasks: gating lifted quality +15.2 points over serving the cheap model alone (95% CI [+12.9, +17.5]) and regressed not one of the 974 — because a rung is only ever spent after a check says the cheap answer is not good enough, and on this workload the gate never once rejected an answer that was already right. See the numbers, including the ones that qualify them.


How it works

Four moves: 1 Route — open the cheapest rung of the ladder. 2 Prove — gate the real output with tests, a schema, or a judge. 3 Escalate — one rung up only on gate failure, budget-capped. 4 Learn — outcomes feed back so the serve threshold self-tunes. Learn loops back to Route.
  1. Route — every request opens on the cheapest rung of your model ladder. No per-prompt classifier picks the model; the cheap model simply takes the first pass.
  2. Prove — a gate checks the actual output: your unit tests, a JSON schema, an LLM judge (maker ≠ checker), or self-consistency. It reads the real answer, not the prompt.
  3. Escalate — only on gate failure: one rung up, budget-capped, with cross-provider failover on a 5xx.
  4. Learn — outcomes feed back via /v1/feedback; the serve threshold self-tunes so the guarantee tracks your live traffic. No policy model to retrain, ever.

Who decides a request needs the expensive model? The gate — from the cheap model's actual answer. Never a classifier guessing from the prompt. Change what "good" means by editing a gate; there's nothing to retrain.

How a test suite becomes a Firstpass gate. The model's reply arrives as JSON on stdin; run-tests.py extracts the code from its Markdown fences, writes it to a temp directory with your test files, and runs your real command. Three outcomes: all tests pass, so the gate returns pass with score 1.0 and the answer is served; some tests fail, so it returns fail with the pass fraction as a graded score and the ladder escalates; or the runner itself is missing, so it returns abstain rather than a fabricated verdict.

Your existing test suite becomes the gate — scripts/gates/run-tests.py wraps pytest, jest or cargo test. → Gates

Live, animated walkthrough → How it works


The proof

The claim no predictive router makes: on 974 real MBPP coding tasks (fail-closed sandbox, real unit-test gates — committed artifact), Firstpass earned a distribution-free bound of ≤10% wrong answers served at 95% confidence — calibrated risk 5.5%, realized served-failure 7.7% at the threshold — while serving 82% of requests from the cheap tier.

On 974 real MBPP tasks with real test gates: served-failure is 5.5% calibrated risk and 7.7% realized, both under the ≤10% distribution-free target line at 95% confidence; 82% of requests are served from the cheap tier, only 18% escalate. Pre-registered α=0.10, δ=0.05.

The bound is a Hoeffding upper confidence bound — valid for any data distribution, no Gaussian assumptions. It's computed from a real run, not assumed. Your savings depend on your workload, which is why every trace records the always-frontier counterfactual: you measure your number instead of trusting ours.

…and whether routing beat simply picking one model

A guarantee is only half the question. The other half is whether any of this beats picking one model and living with it, so the same 974 tasks were replayed under each policy on identical measurements (sonnet ladder · opus ladder).

The headline: the gate never made a task worse. Not one of 974, on either ladder. It escalates on 19% of traffic, 80% of those escalations turn a wrong answer into a right one, and nothing it did cost a task that the cheap model had already got right.

vs always-cheap vs always-top
quality +15.2 points (paired, 95% CI [+12.9, +17.5]) −2.3 points (95% CI [−3.5, −1.1])
tasks ~150 recovered, 0 regressed 28 lost, 6 won
cost +$1.71 total 57% lower per success

That asymmetry is worth being precise about, because it is a property of this workload and not a law. Escalation can only make things worse in one specific way: the gate rejects a cheap answer that was actually correct, and the rung above then gets it wrong. On these 974 tasks the first half never happened — zero correct cheap answers were rejected, matching the 0.0% false-reject rate in the guarantee artifact — so the regression path was never entered. A gate that rejects sloppily on your traffic can regress, which is exactly why the receipts record every verdict and firstpass ope rehearses a policy against your own logs before it enforces anything.

Two honest qualifications, because a number that can't survive scrutiny isn't worth quoting:

  • The saving depends on your ladder, and it is the ladder that moves it. 57% against an opus ceiling, 12% against a sonnet one — 81% of traffic never reaches the top rung, so a pricier ceiling compounds directly.
  • On MBPP those two ceilings are statistically indistinguishable (0.9476 vs 0.9456, a 2-task difference; an independent re-run of the shared cheap rung flipped 80 of 974 outcomes). So part of the 57% is opus being the wrong ceiling for this workload rather than routing beating a well-chosen baseline. The artifacts say so in their headers.
Reproduce it — each command labels itself and states what it costs
cargo run -p firstpass-bench                    # simulation harness (free, self-labeled SIMULATION)
cargo run -p firstpass-bench -- --live          # live benchmark (your key, ~a few $)

# the distribution-free bound on 974 real MBPP tasks (your key + Docker, ~$5):
curl -sLO https://raw.githubusercontent.com/google-research/google-research/master/mbpp/mbpp.jsonl
FIRSTPASS_CODING_DATASET=./mbpp.jsonl \
  cargo run --release -p firstpass-bench -- --coding-live

The harness recomputes the conformal bound from your run's gate/oracle outcomes with the same pre-registered α=0.10, δ=0.05. Result artifacts and provenance rules live in docs/benchmarks/ (methodology + kill criterion).


Firstpass vs. predictive routers

Predictive routers ⚡ Firstpass
Decides by guessing from the prompt proving the real output
A wrong answer ships silently caught by the gate, escalated
Quality guarantee none ≤10% served-failure @ 95%, earned live
Adapts by retraining a policy model self-tuning threshold + edit a gate
Audit trail a dashboard number hash-chained receipt per decision
A policy change deploy and hope rehearsed first: firstpass ope replays your logs with CIs

And the one good idea predictive routers had — starting on the right model — is already inside Firstpass: a learned start-rung bandit picks where the ladder begins, prediction errors cost only latency, and the gate still decides what ships.

Decision-model routers (Jev) vs Firstpass

A new class of "System One" decision models — TypeSafe's Jev — answers one cheap closed-form question per query ($0.042/M input tokens, no text generation) and routes by asking it to pick a model tier directly. Products built on it (jev-router, prismhq/jev-router) then serve that tier's output unverified. Firstpass's thesis is narrower: predict the start, verify what's served — the same cheap signal can pick where the ladder opens (the start-rung bandit already had a slot for exactly this), but the gate still runs on every attempt, so a wrong prediction costs money or latency, never a wrong answer shipped.

The measured gap (simulation, not live). Firstpass's own σ-sweep (cargo run -p firstpass-bench, n=500) puts a Jev-style unverified router's served-failure at 50.6–56.2% across the noise levels tested, against Firstpass's own gated served-failure of 15.8% on the same suite. Full numbers, methodology, and both addenda: ADR 0013.

[escalation.prior] — experimental, default-off. Simulation said STOP; a real-data replay says PROCEED. The original pre-registered sim (σ-sweep, synthetic noise) found the fused arm tying plain Firstpass at σ=0 and losing at higher noise — history, not the current read. A second pre-registration replayed the same prior mechanism on 2,418 real recorded MBPP outcomes across three ladders, using OpenJev (Apache-2.0, DiffusionGemma 26B-A4B, run locally — not TypeSafe's hosted Jev) as the prior source: pooled $/success $0.01126 → $0.01075 (−4.6%), CI of the difference [-0.00074, -0.00031] excludes 0, served-failure held (0.0786 → 0.0778). Verdict: PROCEED. The win is ladder-dependent — −6.3% on haiku→sonnet, −2.4% on haiku→opus, 0% on gpt-4.1-mini→gpt-5.5 (a ~20x price ratio where skipping the cheap rung never pays) — and OpenJev says nothing about hosted Jev's own accuracy. Full numbers and caveats: docs/benchmarks/openjev-prior-replay.md, addendum in ADR 0013. The block stays default-off — the replay used OpenJev, not hosted Jev, so it isn't evidence about that vendor — but ships because it costs nothing while unconfigured:

[escalation.prior]
provider    = "typesafe"                # the only accepted value today
base_url    = "https://api.typesafe.ai" # default; override for self-hosted/mock Jev
model       = "jev-latest"
api_key_env = "TYPESAFE_API_KEY"        # env var name only — the key itself is never logged
timeout_ms  = 150                       # any error fails open — no prior, no `decision_prior` field
strength    = 10                        # Beta pseudo-count weight of the prior
rungs = [
  "rung 0 (claude-haiku) is the least capable tier that fully handles this request",
  "rung 1 (claude-sonnet) is the least capable tier that fully handles this request",
]

rungs needs exactly one entry per rung of every enforce-mode route's ladder — a length mismatch is rejected at config parse, never silently truncated. The gate still verifies every served output; the prior only ever moves where the ladder starts.

Blending in a learned signal: BLEND-NEUTRAL. A second pre-registration (specs/prior-blend-and-decision-gate.md) tested whether blending traffic-learned pass rates into the prior beats the prior alone. Pooled (n=2418): prior+learned $0.01094 vs prior $0.01075 — paired diff +0.00019 [+0.00004, +0.00036], excludes 0, the blend is worse. The prior alone stays the recommendation. (This run also caught and corrected a hindsight leak: an earlier "cost-aware learned-p" arm decided and bucketed on each task's own realized cost, which only exists after generation — its $0.00919 pooled figure is optimistic by $0.00190/success. The honest ex-ante version is $0.01109 pooled, only ~1.5% under first-pass's $0.01126; it's now labeled learned-p (hindsight) in every report and any earlier "~22% cheaper than first-pass" claim for that arm is withdrawn — see ADR 0013.) Full numbers: docs/benchmarks/prior-blend-replay.md.

decision gate — NOT-RECOMMENDED (measured with OpenJev; hosted Jev unmeasured). The same Jev model can also sit behind a [[gate]] decision = {...} block as a cheap external verifier instead of a frontier LLM judge. A transport error, timeout, or unparseable reply ABSTAINs — never a fabricated pass — but scored against 974 served MBPP answers with VRBench's hidden-test oracle (111 oracle-wrong) and local OpenJev as the verifier: catch rate 0.2162 [0.1441, 0.2973], collateral 0.1031 [0.0834, 0.1228], AUC 0.6310 [0.5711, 0.6886] — below the pre-registered bar (catch ≥ 0.30 and collateral ≤ 0.05). Verdict: NOT-RECOMMENDED at τ=0.5. This measures OpenJev, not hosted Jev, which remains unmeasured. A live smoke test (not a measurement) separately caught the gate reading the wrong wire field (probability/p/value instead of the real noul) — fixed. Full numbers: docs/benchmarks/decision-gate-study.md.

[[gate]]
id       = "verify"
decision = { provider = "typesafe", model = "jev-latest", threshold = 0.6 }
# api_key_env defaults to TYPESAFE_API_KEY; base_url defaults to https://api.typesafe.ai

OpenJev (local)

Both [escalation.prior] and the decision gate speak the same POST /v1/systemone contract, so either can point at a locally-run OpenJev (Apache-2.0, DiffusionGemma 26B-A4B — razorback16/openjev) instead of TypeSafe's hosted Jev — this is what docs/benchmarks/openjev-prior-replay.md measured.

# Apple Silicon, ~16 GB unified memory — binds 127.0.0.1:8080
OPENJEV_BACKEND=mlx python -m openjev
# NVIDIA, 24 GB+ VRAM:
# docker compose up
[escalation.prior]
provider    = "typesafe"              # still the only accepted value — OpenJev speaks the same wire contract
base_url    = "http://127.0.0.1:8080"
api_key_env = "TYPESAFE_API_KEY"      # the client always sends a bearer token — set any placeholder, e.g. TYPESAFE_API_KEY=local

api_key_env is required even against a keyless local server — the client always attaches bearer_auth, so a missing env var disables the prior fail-open rather than sending an empty token (crates/firstpass-proxy/src/run.rs, crates/firstpass-proxy/src/gate.rs). Point the decision gate's base_url at the same server to use OpenJev as the verifier instead of the prior.

See docs/related-work.md for how Jev-style routers, RouteLLM, Not Diamond, Martian, Sakana Fugu, FrugalGPT, AutoMix, and CP-Router compare on method, not marketing.


The receipt

🧾 Every decision is a hash-chained trace an auditor can re-derive
{
  "trace_id": "0192f3a1-7c4e-7abc-9d21-4e8b1f0a2c33",
  "prev_hash": "9f2c…a1b7",                          // chains to the prior decision — tamper-evident
  "attempts": [
    { "rung": 0, "model": "anthropic/claude-haiku-4-5", "cost_usd": 0.0007,
      "gates": [{ "gate_id": "cargo-test", "verdict": "fail" }] },   // cheap tried first — gate caught it
    { "rung": 1, "model": "anthropic/claude-sonnet-5", "cost_usd": 0.0121,
      "gates": [{ "gate_id": "cargo-test", "verdict": "pass" }] }    // escalated, proven, served
  ],
  "final": { "served_rung": 1, "total_cost_usd": 0.0128,
             "counterfactual_baseline_usd": 0.0630, "savings_usd": 0.0502 }
}

Downstream outcomes flow back via POST /v1/feedback onto a deferred-verdict side table that never alters the sealed record.

Independently auditable. firstpass export writes the sealed log as JSONL; anyone — an auditor, a regulator, you — runs firstpass verify --file receipts.jsonl on their own machine to re-derive the hash chain from genesis, no proxy and no database in the loop. A single altered or reordered receipt breaks the chain at its index and exits non-zero. Black-box routers can't produce this artifact; it's the EU-AI-Act-style logging story, built in.


Install

No Rust, no toolchain — grab a binary and go:

curl --proto '=https' --tlsv1.2 -LsSf https://github.com/dshakes/firstpass/releases/latest/download/firstpass-proxy-installer.sh | sh

Or through your package manager — each row is live and republishes on every release:

🐍 pip / uvx pip install firstpass · uvx --from firstpass firstpass-proxy
🍺 Homebrew brew install dshakes/tap/firstpass-proxy
🐳 Docker docker run -p 8080:8080 -e FIRSTPASS_BIND=0.0.0.0:8080 ghcr.io/dshakes/firstpass:latest
🦀 Cargo cargo install firstpass-proxy (crates.io, live since v0.4.0; needs a Rust toolchain)
📦 npm npm i -g firstpass-proxy (not firstpass — that name on npm is an unrelated CLI)
⬇️ Binaries macOS · Linux · Windows, checksummed, self-updating (firstpass-proxy-update) — Releases

Quickstart

Three lines. Zero config. Zero risk — observe mode changes nothing:

firstpass-proxy                                     # watches your traffic, touches nothing
export ANTHROPIC_BASE_URL="http://127.0.0.1:8080"   # your agent now routes through firstpass
# … use your agent normally — every call gets a receipt: what it'd route, what you'd save

Convinced by your own numbers? Switch on routing:

cp firstpass.example.toml firstpass.toml
FIRSTPASS_MODE=enforce FIRSTPASS_CONFIG=./firstpass.toml firstpass-proxy

Or skip the env var entirely and let Firstpass start the agent already pointed at it:

firstpass launch claude          # also: codex, or `openai -- <any OpenAI-compatible command>`

It refuses to start if no proxy is listening — or if something that isn't Firstpass is holding the port — because an agent launched at the wrong address fails in a way that reads like the agent is broken.

Leaving is unset ANTHROPIC_BASE_URL. That's the whole offboarding story.

🤖 …or let an agent do it — one command does everything

Don't follow docs. Firstpass detects your machine, plans the setup, executes it, and verifies itself:

$ firstpass onboard --apply
detected: shell=zsh · proxy_running=false · routed=false · claude_cli=true

✓ proxy started (pid 17005, observe mode) — log: firstpass-proxy.log
✓ wired ~/.zshrc — export ANTHROPIC_BASE_URL=http://127.0.0.1:8080
→ optional: claude mcp add firstpass -- firstpass mcp
✓ verified — proxy healthy · capabilities live

Auto-detects your shell (zsh/bash/fish), whether the proxy is running, whether you're already routed, and which agents you have — then does only what's missing. Idempotent (re-run any time), transparent (firstpass onboard alone is a dry run showing the exact plan), and reversible (firstpass offboard strips the shell line, stops the proxy, prints the unset). For agents onboarding themselves: llms.txt + AGENTS.md ship machine-readable setup, GET /v1/capabilities gives runtime discovery, and firstpass mcp exposes traces, savings, evals, policy rehearsal, and receipt verification as tools.

⚡ …or just press the button

Three surfaces, no questions asked in any of them:

$ firstpass kiosk          # finds the provider key already in your environment,
                           # writes a config, puts ONE REAL request through the
                           # ladder, prints the receipt. No key? It runs the
                           # keyless demo instead — the button always works.

$ firstpass handshake      # the same thing for a caller with no terminal: one
                           # JSON document — keys found, the config that would
                           # run, a KEYLESS self-test of route→gate→escalate→
                           # serve→receipt→chain, and the env var to route.
                           # It reports; it never writes. Also an MCP tool.

And once anything is running, GET /panel is a live receipts view served by the binary itself — every rung tried, what the gate said, what it cost against always calling the top of the ladder. No CDN, no build step, no external origin: the receipts stay behind the same tenant boundary as /v1/receipts.


Architecture

Firstpass is a proxy in front of your provider calls. Your agent keeps its existing endpoint — Firstpass speaks both inbound wire dialects and returns a normal response, byte-identical to the served rung.

Architecture: your agent calls the proxy on POST /v1/messages (Anthropic dialect) or POST /v1/chat/completions (OpenAI dialect). Inside, the request flows Route → Gate → Escalate → Serve. Rungs call providers with your own keys — anthropic, openai, or any OpenAI-compatible endpoint — with failover on 5xx. Every decision writes a SHA-256 hash-chained receipt; deferred outcomes on /v1/feedback tune the serve threshold.

Wire APIs

Point anything at it. Firstpass speaks three inbound dialects and translates across vendors, so a cascade can cross providers without the caller knowing:

Endpoint For
POST /v1/messages Anthropic Messages — Claude Code and anything Anthropic-native
POST /v1/chat/completions OpenAI Chat Completions, and every OpenAI-compatible client
POST /v1/responses OpenAI Responses — newer OpenAI-family agents, including tool calls
GET /v1/models Model discovery, so an agent CLI can populate its picker

Crossing vendors is handled, not assumed: a caller that sets reasoning_effort keeps it when the ladder escalates onto an Anthropic rung (as thinking), and a tool call survives the round trip through /v1/responses in both directions rather than being quietly dropped.

Providers

Eight OpenAI-compatible platforms work with no [[provider]] block — name the rung and set the key: groq, deepseek, together, fireworks, mistral, openrouter, xai, cerebras, alongside the built-in anthropic and openai.

[[route]]
mode = "enforce"
ladder = ["groq/llama-3.3-70b-versatile", "anthropic/claude-sonnet-5"]
gates = ["non-empty"]

[[price]]   # required — see below
model = "groq/llama-3.3-70b-versatile"
input_per_mtok = 0.59
output_per_mtok = 0.79

Endpoints ship built in; prices do not, deliberately. A base URL is a stable fact. A price is not — these platforms change pricing without notice, and a stale built-in would write a wrong cost_usd into a tamper-evident receipt and mis-feed your [budget] caps. So a rung on one of these still needs an explicit [[price]], and the proxy refuses to start without it, naming the model and printing the block to paste. firstpass doctor also fails loudly when a configured rung's API key is missing, rather than letting the first real request discover it.

Every provider, including open-source

A ladder rung is <id>/<model> — open on a free local model, escalate to a frontier model only on proven need:

[[provider]]
id = "groq"                                  # any OpenAI-compatible host — Groq, Together,
dialect = "openai"                           # DeepSeek, Mistral, xAI, Azure, an aggregator,
base_url = "https://api.groq.com/openai"     # or your own Ollama / vLLM box
api_key_env = "GROQ_API_KEY"

[[route]]
match  = {}
mode   = "enforce"
ladder = ["groq/llama-3.3-70b-versatile", "anthropic/claude-sonnet-5"]
gates  = ["unit-tests"]

anthropic and openai are built in; Gemini (dialect = "gemini"), AWS Bedrock (auth = "aws_sigv4"), and Google Vertex (auth = "gcp_oauth") use the same shape. Every variant ships in firstpass.example.toml, guarded by a parse test.

Verification status, stated plainly. The Anthropic path is live-verified end-to-end (real traffic through the running proxy). The OpenAI-compatible, Gemini, Bedrock, and Vertex adapters are implemented and offline-tested against recorded wire shapes, pending live verification — each flips to verified when a key-gated CI smoke test exercises it against the real endpoint (roadmap, Phase 1).

Gates — "do I have to write them?"

No. Meet it where you are:

Effort You get
None — observe mode Firstpass reports what it would route and save. Nothing changes.
One sentence — judge gate A second model grades every answer against your plain-English rubric.
One config line — consistency gate The model answers k times; agreement is measured confidence (self-consistency, Wang et al. 2022).
Your existing tests The strongest gate: generated code ships only if your suite actually passes.

Flaky gates auto-disable on an error budget — one bad check can't take down a route.

Long conversations

Agent sessions get long, and that breaks routers in ways that are invisible until you read a bill:

  • A prompt too large for the cheap rung escalates instead of failing the request. It is a 400 from the provider, and treating every 400 as fatal kills a request the next rung up would have served. The receipt records it as context_overflow, so capacity-forced escalations are not pooled with quality failures in your statistics.
  • A session that had to escalate starts there next turn ([escalation.session_promotion]), instead of re-paying for the rung that already failed it — with a periodic downward probe so a promotion is never a one-way ratchet.
  • Cached prompts are billed as cached. Prompt caching splits the prompt across three counters at three different rates; counting only input_tokens reports a 190k-token cached prompt as about 20 tokens. Firstpass prices all three, so the receipt and your [budget] caps see what the call actually cost.
  • A conversation that fits nowhere is condensed rather than refused ([escalation.condense], off by default) — but only once every rung has overflowed, where the choice is a degraded answer versus none at all.

Running more than one replica

The verified cache and session promotion are in-process by default. Behind a load balancer that means each replica keeps its own — the same answer is cached N times over, the hit rate drops roughly by N, and, the part that matters, a retraction only reaches the replica that received the feedback while the others keep serving an answer that has been disproven.

Point both at Redis to share them:

[escalation.verified_cache]
ttl_secs  = 900
redis_url = "redis://cache:6379/0"

[escalation.session_promotion]
after_failures = 2
window         = "30m"
redis_url      = "redis://cache:6379/0"

Requires a build with the redis-cache feature (cargo install firstpass-proxy --features redis-cache). Setting redis_url without it, or pointing it at an unreachable server, fails at startup rather than quietly falling back — a cache that silently stays per-replica looks exactly like one that is working.

Modes

One header, five profiles — set per request via x-firstpass-mode (or per route / env): cost · balanced · quality · latency · max. Same ladder, different serve threshold and escalation appetite: cost serves the cheapest thing that clears the gate, quality/max climb sooner, latency prefers the speculative path.

The science

Firstpass is precise about what's novel versus assembled from known parts (the cascade itself is prior art):

  • Learned start-rung bandit — deterministic UCB1, or Thompson sampling with discounted Beta posteriors, drift-forgetting, and logged native-MC propensities (ADR 0007). Predicts where to start; the gate still decides what to serve.
  • The guarantee — split-conformal (Hoeffding UCB) or Learn-then-Test / RCPS exact-binomial testing (firstpass calibrate --method ltt), tracked live under drift by adaptive conformal (Gibbs–Candès). Two Prometheus gauges expose the loop: firstpass_serve_threshold, firstpass_realized_served_failure.
  • Off-policy evaluation — firstpass ope replays your logged receipts against a candidate ladder with IPS / SNIPS / DR estimators and confidence intervals: rehearse a policy change before you ship it.
  • Elastic verification (validated research, phase-1 shipped, not default-on) — a cheap k-sample probe decides how much proof to spend: unanimous-wrong escalates immediately, unanimous-right serves without the expensive gate, only the uncertain middle pays for it (ADR 0008).
Elastic verification: a cheap 5-sample probe decides how much verification to spend. Agreement 0/5 → escalate now (12% of traffic, 0% oracle-correct). Agreement 1–4/5 → run the full gate (23%, mixed). Agreement 5/5 → serve without the expensive gate (65%, 99% oracle-correct). 77% of traffic is decided by the cheap probe alone. Validated, k=5, n=150.
⚙️ Configuration — 12-factor, env-driven
Variable Purpose Default
FIRSTPASS_MODE observe | enforce observe
FIRSTPASS_BIND listen address 127.0.0.1:8080
FIRSTPASS_CONFIG path to firstpass.toml (routes, ladders, gates, providers) —
FIRSTPASS_DB trace store path firstpass.db
FIRSTPASS_RECEIPTS best_effort | durable — durable spills receipts to disk under backpressure instead of dropping, and drains them on boot (audit chain stays valid) best_effort

Endpoints: POST /v1/messages (Anthropic drop-in) · POST /v1/chat/completions (OpenAI drop-in) · POST /v1/feedback · GET /v1/capabilities · GET /healthz · GET /metrics.

Multi-tenant deployments add per-tenant auth (Argon2id), rate limits, gate-health scoping, and AES-256-GCM key custody — all opt-in, default-off (ADR 0004).


Status

v0.3.0 — pre-GA, shipped in the open. Honest about the line between shipped and researched.

✅ Shipped & verified 🔬 Next / research
Both wire dialects, structured enforce default-on Elastic verification (validated, phasing in)
All six gate kinds + per-gate on_abstain (incl. the experimental, unmeasured decision gate) Cross-dialect structured translation beyond Anthropic↔OpenAI
Start-rung bandit (UCB1 / Thompson), speculation, failover Four provider dialects await live wire verification
Conformal guarantee + Learn-then-Test 30-day soak, external security audit
Adaptive threshold, OPE, savings / evals Hosted multi-tenant plane
Receipts + export/verify + durable mode crates.io publish
Modes, per-deployment [[price]], Grafana dashboard, nightly provider-smoke CI

GA is a checklist we publish (ADR 0003), not an adjective — the exact remaining items (secrets, soak clock, external audit) are enumerated in the GA handoff.


Links

Docs · How it works · The guarantee · SPEC · Example config · ADRs · Agent guide · llms.txt · License

Try cheap. Prove it. Escalate only on failure.

proof over prediction · receipts over adjectives

PRs here are gated by compass: agent review · security · cross-model audits · tests — then a human merges.

About

The verification layer for LLM serving — your check runs on every answer before it ships, with a signed, tamper-evident receipt for each decision and a distribution-free bound on wrong answers served. Proof over prediction.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages