Skip to content

feat: provider-side tools — a model declares hosted web search and the transcript shows it (stage 33) - #41

Merged
benjipeng merged 21 commits into
mainfrom
feat/provider-side-tools
Sep 22, 2026
Merged

benjipeng merged 21 commits into
mainfrom
feat/provider-side-tools

Conversation

@benjipeng

@benjipeng benjipeng commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

What

A model may declare the hosted tools its route accepts (server_tools = ["web_search"], PRV-6). The harness forwards the declaration in each dialect's spelling, decodes what the provider returns, and shows a provider-run search as a transcript row in its own colour: [~] web_search · running where the provider began it, finishing once as [+] web_search · <query> or [!] … · failed (PRV-5, ENT-2). The harness never runs the search. Reopening a session shows the same rows (ENT-3).

Why this shape

Decided in .agents/spikes/provider-side-tools/ against Codex, pi, grok-build and claude-code, and measured on one live gateway route. Two findings that shaped the code: the route delivers every done after the last message, so the row is placed at added; and the route delivers a whole turn in one burst after ~30 s, which is upstream of both the gateway and this app (recorded in the spike, not worked around).

Verified

  • Local: workspace tests 1614 passed (unsandboxed, TCP-binding CLI tests included), clippy -D warnings, fmt, the seven pre-commit gates, scripts/smoke-server-tool.py at three widths plus reopen, offline replay of the captured gateway stream through the real binary.
  • CI on 8cdd911: Rust and terminal tests, static and script checks, and macOS Apple Silicon all pass.
  • Not observed: a completed live TUI run on the owner's gateway beyond the captured stream.

Also on this branch

feat(provider): let a route keep its bearer token in the file (8cdd911) — an inline api_key on a route, with a set environment variable still overriding it, written by another agent at the owner's request. Committed on its own after the pre-commit gates were run by hand; CI covers it with the rest.

…ses carry it

A provider that runs a search on its own side is the category the fence position
does not cover: not something we execute, not something we hand to a subprocess,
but something the provider performs inside a call the owner already authorised.
PRV-5 records that such a block is refused today. It does not say whether that
should stand, and this is the evidence for deciding.

Measured against the local gateway rather than read from vendor documentation.
Muse Spark turns out to be a Meta model speaking the Messages protocol, so its
`owned_by: anthropic` names the wire format and not the vendor — a test resting
on that route proves an adapter speaks Messages, never what Anthropic's own
service does. Its effort ladder accepts four of the six levels it declares, and
the OpenAI facade rewrites `reasoning_effort` into the Messages field, so both
routes land on one parameter.

Both dialects offer one hosted search tool and no fetch tool: opening a page is
an action inside search. Responses reports the queries as a `web_search_call`
item, Messages as `server_tool_use` blocks, and neither reports the findings.
Codex reaches the same place from the other vendor, setting `results: None` on
its inline path, so the hole is the shape of inline hosted search rather than a
local defect. Replay has no structural blocker here: this route accepts the
assistant turn back with those blocks present or stripped.

The comparison found two shapes that are not variants of each other. Codex makes
it a first-class protocol item, paying for it at every site that already matches
its response enum. pi makes it an ordinary client tool whose execution issues an
isolated request, so the main conversation never carries a server block at all.

pi's gating is rejected and the reason is already an invariant here: PRV-6 says
configuration names data and selection never uses fuzzy names, while pi matches
provider names and `id.startsWith("claude-")`. A declared `server_tools` subset
follows `allowed_reasoning_efforts` instead — config names the capability, the
adapter owns its dialect spelling, and an unencodable declaration fails before
network work. Today's `max` correction is that failure mode working as intended.

Evidence is reproducible: the gateway commands are in the spike, keyed from the
environment rather than inline.
…g built

The spike found two shapes that are not variants of each other, and left the
choice open. This closes it and orders the work.

The harness forwards a declaration and renders what comes back. It does not run
the search, does not infer the capability from a model's name, and does not
silently absorb a block it cannot show. pi's isolated client tool is recorded as
rejected with the reason that only appears against this codebase: a tool issuing
its own model request needs to reach a provider route, which is a larger
boundary change than the one it saves, and it spends a second model call per
search that the owner never selected.

The declaration follows `allowed_reasoning_efforts` rather than inventing a
pattern: present or absent, nonempty and unique when present, encodable by the
model's dialect, and failing at config load rather than at run time. PRV-6
already refuses authority in a name, so gating on `id.startsWith(…)` was never
available here even though every harness read does it that way. Config names the
capability and the adapter owns its dialect spelling, which is where the
compatibility work belongs.

Six slices in an order with a reason: a spelling cannot be tested with nothing
to spell, a fixture is cheaper to trust when we made the request that produced
it, a row has nothing to render until an event exists, and replay picks its arm
knowing what the row had to show. The transcript slice carries three widths
because the contract owns the row and string assertions do not prove it.

Two questions stay open inside the stage rather than being guessed now: whether
a provider-side row reuses the tool row with a marker naming the actor, which
rendered frames decide, and where hosted-tool spend belongs, since the usage
counter names requests rather than tokens.

PRV-5 currently states that a provider-side block is refused. It was true when
written and this stage rewrites it, which the plan records as deliberate so the
corpus does not carry two readings.

Spike sections that argued for the shape we did not take are rewritten to match,
so the spike and the plan do not disagree in front of the next reader.
…laration and discards thinking

Muse over /v1/chat/completions, streaming, 4096 tokens, one prompt asking for
the latest stable Rust release. Without tools it answers 1.98.0 from memory in
1733 tokens. With the Anthropic-shaped hosted tool it answers 1.98.0 in 2684;
with the OpenAI-shaped one, 1.98.0 in 2163. Every run returns 200, no
tool_calls delta and no field beyond role and content. The same prompt over
/v1/messages with the hosted tool searched four times and answered 1.98.1, a
release published after the model's knowledge.

So a hosted-tool declaration on the Chat Completions route is swallowed with
no error, which is the evidence behind slice 1: the declaration has to fail at
config load, because the wire will not say.

The same route discards reasoning entirely. A one-word answer at high effort
costs 257 completion tokens and arrives as role and content only, in both
streaming and non-streaming form. Over Messages the same model returns
redacted_thinking blocks and a thinking_tokens count, so Muse's reasoning is
reachable through Messages as opaque replay and through Chat Completions not
at all.
…name

Slice 2 claimed Chat Completions and GenerateContent cannot encode a hosted
tool. Both can: OpenAI's own search models take web_search_options, and Gemini
has a google_search tool. So slice 1's gate is not the dialect but the name, and
a declared tool a dialect has no spelling for is what fails at config load.

What the local gateway does with the Chat spelling is a separate fact, now
recorded with its mechanism. At upstream snapshot v7.3.9 the request translator
passes only function tools toward Claude and skips every other type without an
error, and nothing reads web_search_options. The response translator has
branches for tool_use and thinking and none for server_tool_use or
redacted_thinking, so both fall through. The second hole has no OpenAI-standard
shape to fill it with, because url_citation needs URLs the route never returns.

A spelling the route ignores is the owner's declaration being wrong, and on this
facade the wire will not say so. The plan now says a declaration is verified
once rather than trusted.
… has a native Meta route

Two earlier readings were wrong and are rewritten in place.

First, hosted search does not return queries and never findings. That was the
Messages surface as observed through the gateway, and OpenAI's inline path as
Codex models it. Meta's own Responses documentation returns url_citation
annotations with url, title and span on the answer text, and on request the
results themselves with title, url and snippet. The plan's transcript slice and
its exclusion list no longer claim that no dialect returns findings; the row
shows what its route returned.

Second, the Chat Completions holes are not the gateway translator's alone.
Meta's Chat Completions has no search grounding and does not carry reasoning
across turns, in its own words. The translation layer removes what the protocol
would carry, and the protocol carries neither of the two things this stage
needs. So moving Muse to an OpenAI-compatible upstream to keep Chat Completions
native would change nothing that matters.

cli-proxy-api declares meta-api-key as a provider kind. MetaKey aliases
CodexKey, its executor speaks Responses to api.meta.ai/v1, and it forwards
tools, include and store untouched while deleting five named fields and
stripping orphan reasoning ids when store is false. The stack routes Muse
through claude-api-key today, so every surface but Messages is a translation.

Plexmaton's Responses encoder already sends store false and asks for encrypted
reasoning, and its decoder already retains annotations as replay metadata. A
web_search_call item is a typed refusal on both added and done, which is the
PRV-5 refusal this stage rewrites.

Every Meta Responses claim is marked documented rather than observed: a direct
probe was refused by the permission classifier as credential extraction, and
the gateway has no native Meta route configured to probe through.
…route is decided

Through the gateway's existing route, still translated to Messages and back, a
Responses request with tools web_search returns web_search_call items carrying
search actions, reasoning items carrying encrypted_content, and an answer only
a search could give. No gateway change is needed for a Session on
openai_responses to reach hosted search on this model.

The empty control run was a budget exhausted while thinking, returned as
status incomplete with reason max_output_tokens, which is the typed state PRV-5
already maps. Not a route defect.

The translated route loses three things, each observed: annotations are empty
and no results arrive with or without the include; open_page actions arrive as
search with an empty query, so a decoder accepts an empty query; and
reasoning_tokens reads zero. The gateway's native meta-api-key route would
restore all three and is the owner's to switch to later. It changes nothing
about which items a decoder has to accept.

The plan now decides Responses and defers Messages: server_tool_use keeps its
PRV-5 refusal until a Messages route is in daily use, and the one semantic event
is shaped so a Messages decoder can emit it later without a second one.
Five sentences described the local gateway rather than what the stage builds:
that Chat Completions swallows a declaration without saying so, that a route
flattened an action into an empty query, that the owner's model reaches search
over Responses, that this gateway refuses three other hosted tools, and that a
native meta-api-key route would restore citations. Each is true of one host and
none is a property of the code. They are now the general rule they stood for: a
declaration is the owner's claim about their route, an action without a query is
carried rather than refused, the Messages decoder waits for a route in daily
use, no measured route offered a second hosted tool, and the row shows what its
route returned. The spike keeps the measurements, which is what a spike is for.

The phase entry claimed both dialects return queries and neither returns
findings. That was one route's behaviour and is no longer believed of the
protocol; the entry now says the evidence was measured against one live route.
The replay slice leaned on what one gateway accepts. It now states the rule the
stage builds to: a hosted-tool item round-trips under PRV-3's sidecar rules, its
provider id need not be preserved, and sending it back or stripping it is a
decision the test names.
…ough the executable

Two turns through main's binary with muse chosen via /model: an arithmetic
answer, then a follow-up that depends on it, so the first turn replayed through
the gateway's translation and back. Both correct, usage reported, nothing on
screen read as an error. Whether the encrypted reasoning survived the replay is
not observable from the screen and is recorded as untested.
Slice 1 of stage 33. A model entry gains server_tools, a subset of the
capabilities a provider can run on its own side inside the call the owner
already authorised. Today that catalog has one name, web_search.

The field follows allowed_reasoning_efforts rather than inventing a pattern:
absent means none, present means nonempty and unique, and every name must be
one the model's dialect can spell. Responses, Chat Completions and Messages
spell web search; GenerateContent has a google_search tool no route here has
been measured with, so a declaration on it fails at config load rather than
being guessed at. The spelling gate lives on ModelApi so the encoders slice 2
adds consult the same answer the validator did.

The gate runs at config load, before any network work, because the wire will
not always say: a route can accept a spelling it does not understand and ignore
it. PRV-6 now states the declaration, that nothing infers it from a provider or
model name, and that a route ignoring a correct spelling is the declaration
being wrong for the owner to fix.

A first draft bounded the list's length by the catalog size as well as by
uniqueness. With one catalog entry the two checks are the same check, and
mutating the uniqueness test away could not fail a test. Uniqueness over a
finite enum already implies the bound, so the bound is gone and every remaining
branch has a mutation that kills its test.

Verified locally: plexmaton-provider 54 config tests pass (three new, each
proven by mutation), clippy clean with -D warnings, fmt clean, citations, doc
budget, file length and typos clean. Not run: the workspace suite, CI.
Slice 2 of stage 33. Configuration named the capability in slice 1; the
dialect now owns its spelling and nothing else does. Responses appends
{"type":"web_search"} to the tools array, Messages appends the dated
web_search_20250305 server tool, and Chat Completions sets web_search_options,
which is where that schema puts hosted search rather than among the tool
types. Hosted tools follow the function tools, so the order an owner reads in
the body is the order declared, and a hosted tool alone still names a tool
choice. GenerateContent asserts no declaration reached it: PRV-6 refuses one at
config load, and stating the invariant at the encoder keeps a future spelling
from landing on one side and not the other.

One test asserts the emitted JSON for all three dialects, with a function tool
beside the hosted one and without, and that an undeclared model carries
nothing. Removing any one spelling fails it.

Verified locally: provider crate passes (41 fixture tests, one new), the three
spelling mutations each fail the test, clippy clean with -D warnings, fmt,
citations, doc budget, file length and typos clean. Not run: the workspace
suite, CI.
… item it was

Slices 3 and 5 of stage 33. A Responses web_search_call is no longer a typed
refusal. It becomes a server-tool call in the record: what the provider did, a
search, an opened page or a find within one, bounded the way tool arguments are
and counted with them; the provider's own item identity and status ride beside
it as replay metadata, the way a function call's do. It reaches no admission
and no scheduler, because the call already happened inside the model call the
owner authorised by declaring the route accepts it.

The vocabulary lives in core beside ReasoningEffort, since configuration,
the record and, next, the transcript all name it; the provider re-exports
ServerTool so slice 1's declaration is unchanged. The event is a sixth
ModelEvent and the block a fifth AssistantBlock, and every exhaustive site
says what it does with them: the step records it without a call identity of
ours, the compaction collector keeps it beside the summary text, degradation
drops it for a foreign model the way it drops replay-only blocks because the
answer text carries what came of it, the projection claims its entry and
projects nothing yet, Chat refuses it as unrepresentable, and the Messages
decoder refuses to emit one until a Messages route is in daily use.

Decoding follows what live routes send. Both spellings of a query arrive,
sometimes together, so one is kept once; a search that arrived with no query
is carried as it arrived, because one route flattens an opened page into
exactly that; an action kind the record cannot name fails the step rather than
travelling blind, which is where this differs from Codex's catch-all and why
PRV-5 records that choice as rejected. The progress markers a route streams
while searching carry nothing the finished item lacks and are ignored.

On replay the Responses encoder rebuilds the item from the record and adds
identity and status from the sidecar, so the next request reads as the
provider wrote it, in the order it wrote it. The fixture is sanitized from a
live stream through the local gateway, with one action changed to an opened
page so both observed kinds are exercised, and the same test drives a
follow-up turn to prove the replay.

PRV-5 is rewritten in this change rather than left as a second reading, and
PRV-3 names the new identity it retains. The plan's slice 5 is implemented
here because the encoder's exhaustive match had to gain the arm anyway, and
an arm that could not be proven would have been the worse choice.

Verified locally: core, agent and provider suites pass with five new tests,
the runtime suite passes outside the sandbox, clippy clean with -D warnings on
all four crates, fmt, citations, crate graph, doc budget, file length and
typos clean. Removing the decoder arm or the replayed identity each fails the
fixture test. Not run: the TUI and CLI suites, the PTY smokes, CI.
…le history

The previous commit added an encode error the status snapshot's exhaustive
match did not cover, so the workspace stopped building. Only four crates were
compiled before that commit; the site count had covered the block and event
enums and not this one. The variant belongs with opaque replay in Chat and
unrepresentable order: history this dialect cannot carry. The content-free
category test gains the row that proves it.

Verified locally: the whole workspace builds with all targets, the full suite
passes with the fence gate on, clippy is clean with -D warnings on every crate.
…rovider id

Driving this branch's binary through a search turn and a follow-up found the
replay shape wrong for the route in daily use. The first turn searched and
answered 1.98.1; the follow-up, which replayed the recorded web_search_call
items, came back "provider returned HTTP 400". Bisected against the gateway:
any replayed item carrying an id fails with "content block
web_search_tool_result is not supported on assistant messages", because the
gateway's Responses-to-Claude translator rebuilds a search-and-result pair for
an id it recognises and the upstream refuses the result half on an assistant
message. The same item without an id is accepted and answered. Meta documents
the id as optional on replay.

So the encoder replays the item from the record with the sidecar's status and
keeps the id in the sidecar only. PRV-3 records the rule and rejects the
alternative, which is the shape Codex sends and the one that fails here. The
fixture test asserts the id is absent, and the live follow-up now answers.
…with and without its id

Read after the bisect rather than inferred from the error: the translator
rebuilds an id-bearing item as a server_tool_use and web_search_tool_result
pair on the assistant message, which the upstream refuses, and drops an
id-less item entirely before Meta sees it. The second half is what makes
replaying without the id harmless on this route, and what a native route
would do differently.
Stage 33 slice 4. A server tool call now reaches the workspace as one
event in its terminal state and one row in the tool row's grammar:
`[+] web_search · <what the route reported>`, disclosure and copy of
the same facts, counted among the tools. The marker wears a new
seventeenth role spent on the palette's purple slot, so who ran the
call is a colour and what it was is the name; a call the provider
gave up wears failure like any other, which needed the outcome typed
into the record. Ending states the decoder did not know (failed,
searching) are now read rather than failing deserialization.

The user chose the grammar and the shade from rendered candidates and
reviewed the real frames at three widths on 2026-09-21; the purple and
magenta slots take the values they picked. ui-ux, ENT-2, PRV-5 and the
plan say what the code now does.
A live gateway finishes every web_search_call after the last message,
so a row placed at `done` landed after the answer it informed, and a
reopened conversation, which orders by position, showed the rows
somewhere else. The Responses decoder now reports the call when its
item is added; the step places it as a running row at that position
and finishes it at the next revision, or reports it failed if the step
ends first, keeping no block for a call that never finished. The
journal holds only the finished call, so replay shows the same row at
revision zero, in the same place.

The activity line names the running search like a local tool's. A
loopback smoke drives the burst shape the gateway produced through the
binary and reopens the conversation to prove the order holds. ENT-2,
PRV-5, ui-ux and the plan say what the code now does; the spike
records the measurement that forced the change.
…ateway does

The loopback burst now carries the reasoning items a live route sends
before each message, so the smoke exercises replay beside the placed
and finished server tool rows, which is the shape a captured gateway
stream renders correctly once the terminal is actually read.
…he environment still overriding

Gates run by hand before this commit (fmt, file-length, crate-graph, citations, frames, typos: all clean);
the hook's mktemp is unavailable in this shell.
@benjipeng
benjipeng merged commit f007953 into main Sep 22, 2026
3 checks passed
@benjipeng
benjipeng deleted the feat/provider-side-tools branch September 22, 2026 17:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant