feat: provider-side tools — a model declares hosted web search and the transcript shows it (stage 33) - #41
Merged
Merged
Conversation
…ses carry it
A provider that runs a search on its own side is the category the fence position
does not cover: not something we execute, not something we hand to a subprocess,
but something the provider performs inside a call the owner already authorised.
PRV-5 records that such a block is refused today. It does not say whether that
should stand, and this is the evidence for deciding.
Measured against the local gateway rather than read from vendor documentation.
Muse Spark turns out to be a Meta model speaking the Messages protocol, so its
`owned_by: anthropic` names the wire format and not the vendor — a test resting
on that route proves an adapter speaks Messages, never what Anthropic's own
service does. Its effort ladder accepts four of the six levels it declares, and
the OpenAI facade rewrites `reasoning_effort` into the Messages field, so both
routes land on one parameter.
Both dialects offer one hosted search tool and no fetch tool: opening a page is
an action inside search. Responses reports the queries as a `web_search_call`
item, Messages as `server_tool_use` blocks, and neither reports the findings.
Codex reaches the same place from the other vendor, setting `results: None` on
its inline path, so the hole is the shape of inline hosted search rather than a
local defect. Replay has no structural blocker here: this route accepts the
assistant turn back with those blocks present or stripped.
The comparison found two shapes that are not variants of each other. Codex makes
it a first-class protocol item, paying for it at every site that already matches
its response enum. pi makes it an ordinary client tool whose execution issues an
isolated request, so the main conversation never carries a server block at all.
pi's gating is rejected and the reason is already an invariant here: PRV-6 says
configuration names data and selection never uses fuzzy names, while pi matches
provider names and `id.startsWith("claude-")`. A declared `server_tools` subset
follows `allowed_reasoning_efforts` instead — config names the capability, the
adapter owns its dialect spelling, and an unencodable declaration fails before
network work. Today's `max` correction is that failure mode working as intended.
Evidence is reproducible: the gateway commands are in the spike, keyed from the
environment rather than inline.
…g built The spike found two shapes that are not variants of each other, and left the choice open. This closes it and orders the work. The harness forwards a declaration and renders what comes back. It does not run the search, does not infer the capability from a model's name, and does not silently absorb a block it cannot show. pi's isolated client tool is recorded as rejected with the reason that only appears against this codebase: a tool issuing its own model request needs to reach a provider route, which is a larger boundary change than the one it saves, and it spends a second model call per search that the owner never selected. The declaration follows `allowed_reasoning_efforts` rather than inventing a pattern: present or absent, nonempty and unique when present, encodable by the model's dialect, and failing at config load rather than at run time. PRV-6 already refuses authority in a name, so gating on `id.startsWith(…)` was never available here even though every harness read does it that way. Config names the capability and the adapter owns its dialect spelling, which is where the compatibility work belongs. Six slices in an order with a reason: a spelling cannot be tested with nothing to spell, a fixture is cheaper to trust when we made the request that produced it, a row has nothing to render until an event exists, and replay picks its arm knowing what the row had to show. The transcript slice carries three widths because the contract owns the row and string assertions do not prove it. Two questions stay open inside the stage rather than being guessed now: whether a provider-side row reuses the tool row with a marker naming the actor, which rendered frames decide, and where hosted-tool spend belongs, since the usage counter names requests rather than tokens. PRV-5 currently states that a provider-side block is refused. It was true when written and this stage rewrites it, which the plan records as deliberate so the corpus does not carry two readings. Spike sections that argued for the shape we did not take are rewritten to match, so the spike and the plan do not disagree in front of the next reader.
…laration and discards thinking Muse over /v1/chat/completions, streaming, 4096 tokens, one prompt asking for the latest stable Rust release. Without tools it answers 1.98.0 from memory in 1733 tokens. With the Anthropic-shaped hosted tool it answers 1.98.0 in 2684; with the OpenAI-shaped one, 1.98.0 in 2163. Every run returns 200, no tool_calls delta and no field beyond role and content. The same prompt over /v1/messages with the hosted tool searched four times and answered 1.98.1, a release published after the model's knowledge. So a hosted-tool declaration on the Chat Completions route is swallowed with no error, which is the evidence behind slice 1: the declaration has to fail at config load, because the wire will not say. The same route discards reasoning entirely. A one-word answer at high effort costs 257 completion tokens and arrives as role and content only, in both streaming and non-streaming form. Over Messages the same model returns redacted_thinking blocks and a thinking_tokens count, so Muse's reasoning is reachable through Messages as opaque replay and through Chat Completions not at all.
…name Slice 2 claimed Chat Completions and GenerateContent cannot encode a hosted tool. Both can: OpenAI's own search models take web_search_options, and Gemini has a google_search tool. So slice 1's gate is not the dialect but the name, and a declared tool a dialect has no spelling for is what fails at config load. What the local gateway does with the Chat spelling is a separate fact, now recorded with its mechanism. At upstream snapshot v7.3.9 the request translator passes only function tools toward Claude and skips every other type without an error, and nothing reads web_search_options. The response translator has branches for tool_use and thinking and none for server_tool_use or redacted_thinking, so both fall through. The second hole has no OpenAI-standard shape to fill it with, because url_citation needs URLs the route never returns. A spelling the route ignores is the owner's declaration being wrong, and on this facade the wire will not say so. The plan now says a declaration is verified once rather than trusted.
… has a native Meta route Two earlier readings were wrong and are rewritten in place. First, hosted search does not return queries and never findings. That was the Messages surface as observed through the gateway, and OpenAI's inline path as Codex models it. Meta's own Responses documentation returns url_citation annotations with url, title and span on the answer text, and on request the results themselves with title, url and snippet. The plan's transcript slice and its exclusion list no longer claim that no dialect returns findings; the row shows what its route returned. Second, the Chat Completions holes are not the gateway translator's alone. Meta's Chat Completions has no search grounding and does not carry reasoning across turns, in its own words. The translation layer removes what the protocol would carry, and the protocol carries neither of the two things this stage needs. So moving Muse to an OpenAI-compatible upstream to keep Chat Completions native would change nothing that matters. cli-proxy-api declares meta-api-key as a provider kind. MetaKey aliases CodexKey, its executor speaks Responses to api.meta.ai/v1, and it forwards tools, include and store untouched while deleting five named fields and stripping orphan reasoning ids when store is false. The stack routes Muse through claude-api-key today, so every surface but Messages is a translation. Plexmaton's Responses encoder already sends store false and asks for encrypted reasoning, and its decoder already retains annotations as replay metadata. A web_search_call item is a typed refusal on both added and done, which is the PRV-5 refusal this stage rewrites. Every Meta Responses claim is marked documented rather than observed: a direct probe was refused by the permission classifier as credential extraction, and the gateway has no native Meta route configured to probe through.
…route is decided Through the gateway's existing route, still translated to Messages and back, a Responses request with tools web_search returns web_search_call items carrying search actions, reasoning items carrying encrypted_content, and an answer only a search could give. No gateway change is needed for a Session on openai_responses to reach hosted search on this model. The empty control run was a budget exhausted while thinking, returned as status incomplete with reason max_output_tokens, which is the typed state PRV-5 already maps. Not a route defect. The translated route loses three things, each observed: annotations are empty and no results arrive with or without the include; open_page actions arrive as search with an empty query, so a decoder accepts an empty query; and reasoning_tokens reads zero. The gateway's native meta-api-key route would restore all three and is the owner's to switch to later. It changes nothing about which items a decoder has to accept. The plan now decides Responses and defers Messages: server_tool_use keeps its PRV-5 refusal until a Messages route is in daily use, and the one semantic event is shaped so a Messages decoder can emit it later without a second one.
Five sentences described the local gateway rather than what the stage builds: that Chat Completions swallows a declaration without saying so, that a route flattened an action into an empty query, that the owner's model reaches search over Responses, that this gateway refuses three other hosted tools, and that a native meta-api-key route would restore citations. Each is true of one host and none is a property of the code. They are now the general rule they stood for: a declaration is the owner's claim about their route, an action without a query is carried rather than refused, the Messages decoder waits for a route in daily use, no measured route offered a second hosted tool, and the row shows what its route returned. The spike keeps the measurements, which is what a spike is for. The phase entry claimed both dialects return queries and neither returns findings. That was one route's behaviour and is no longer believed of the protocol; the entry now says the evidence was measured against one live route.
The replay slice leaned on what one gateway accepts. It now states the rule the stage builds to: a hosted-tool item round-trips under PRV-3's sidecar rules, its provider id need not be preserved, and sending it back or stripping it is a decision the test names.
…ough the executable Two turns through main's binary with muse chosen via /model: an arithmetic answer, then a follow-up that depends on it, so the first turn replayed through the gateway's translation and back. Both correct, usage reported, nothing on screen read as an error. Whether the encrypted reasoning survived the replay is not observable from the screen and is recorded as untested.
Slice 1 of stage 33. A model entry gains server_tools, a subset of the capabilities a provider can run on its own side inside the call the owner already authorised. Today that catalog has one name, web_search. The field follows allowed_reasoning_efforts rather than inventing a pattern: absent means none, present means nonempty and unique, and every name must be one the model's dialect can spell. Responses, Chat Completions and Messages spell web search; GenerateContent has a google_search tool no route here has been measured with, so a declaration on it fails at config load rather than being guessed at. The spelling gate lives on ModelApi so the encoders slice 2 adds consult the same answer the validator did. The gate runs at config load, before any network work, because the wire will not always say: a route can accept a spelling it does not understand and ignore it. PRV-6 now states the declaration, that nothing infers it from a provider or model name, and that a route ignoring a correct spelling is the declaration being wrong for the owner to fix. A first draft bounded the list's length by the catalog size as well as by uniqueness. With one catalog entry the two checks are the same check, and mutating the uniqueness test away could not fail a test. Uniqueness over a finite enum already implies the bound, so the bound is gone and every remaining branch has a mutation that kills its test. Verified locally: plexmaton-provider 54 config tests pass (three new, each proven by mutation), clippy clean with -D warnings, fmt clean, citations, doc budget, file length and typos clean. Not run: the workspace suite, CI.
Slice 2 of stage 33. Configuration named the capability in slice 1; the
dialect now owns its spelling and nothing else does. Responses appends
{"type":"web_search"} to the tools array, Messages appends the dated
web_search_20250305 server tool, and Chat Completions sets web_search_options,
which is where that schema puts hosted search rather than among the tool
types. Hosted tools follow the function tools, so the order an owner reads in
the body is the order declared, and a hosted tool alone still names a tool
choice. GenerateContent asserts no declaration reached it: PRV-6 refuses one at
config load, and stating the invariant at the encoder keeps a future spelling
from landing on one side and not the other.
One test asserts the emitted JSON for all three dialects, with a function tool
beside the hosted one and without, and that an undeclared model carries
nothing. Removing any one spelling fails it.
Verified locally: provider crate passes (41 fixture tests, one new), the three
spelling mutations each fail the test, clippy clean with -D warnings, fmt,
citations, doc budget, file length and typos clean. Not run: the workspace
suite, CI.
… item it was Slices 3 and 5 of stage 33. A Responses web_search_call is no longer a typed refusal. It becomes a server-tool call in the record: what the provider did, a search, an opened page or a find within one, bounded the way tool arguments are and counted with them; the provider's own item identity and status ride beside it as replay metadata, the way a function call's do. It reaches no admission and no scheduler, because the call already happened inside the model call the owner authorised by declaring the route accepts it. The vocabulary lives in core beside ReasoningEffort, since configuration, the record and, next, the transcript all name it; the provider re-exports ServerTool so slice 1's declaration is unchanged. The event is a sixth ModelEvent and the block a fifth AssistantBlock, and every exhaustive site says what it does with them: the step records it without a call identity of ours, the compaction collector keeps it beside the summary text, degradation drops it for a foreign model the way it drops replay-only blocks because the answer text carries what came of it, the projection claims its entry and projects nothing yet, Chat refuses it as unrepresentable, and the Messages decoder refuses to emit one until a Messages route is in daily use. Decoding follows what live routes send. Both spellings of a query arrive, sometimes together, so one is kept once; a search that arrived with no query is carried as it arrived, because one route flattens an opened page into exactly that; an action kind the record cannot name fails the step rather than travelling blind, which is where this differs from Codex's catch-all and why PRV-5 records that choice as rejected. The progress markers a route streams while searching carry nothing the finished item lacks and are ignored. On replay the Responses encoder rebuilds the item from the record and adds identity and status from the sidecar, so the next request reads as the provider wrote it, in the order it wrote it. The fixture is sanitized from a live stream through the local gateway, with one action changed to an opened page so both observed kinds are exercised, and the same test drives a follow-up turn to prove the replay. PRV-5 is rewritten in this change rather than left as a second reading, and PRV-3 names the new identity it retains. The plan's slice 5 is implemented here because the encoder's exhaustive match had to gain the arm anyway, and an arm that could not be proven would have been the worse choice. Verified locally: core, agent and provider suites pass with five new tests, the runtime suite passes outside the sandbox, clippy clean with -D warnings on all four crates, fmt, citations, crate graph, doc budget, file length and typos clean. Removing the decoder arm or the replayed identity each fails the fixture test. Not run: the TUI and CLI suites, the PTY smokes, CI.
…le history The previous commit added an encode error the status snapshot's exhaustive match did not cover, so the workspace stopped building. Only four crates were compiled before that commit; the site count had covered the block and event enums and not this one. The variant belongs with opaque replay in Chat and unrepresentable order: history this dialect cannot carry. The content-free category test gains the row that proves it. Verified locally: the whole workspace builds with all targets, the full suite passes with the fence gate on, clippy is clean with -D warnings on every crate.
…rovider id Driving this branch's binary through a search turn and a follow-up found the replay shape wrong for the route in daily use. The first turn searched and answered 1.98.1; the follow-up, which replayed the recorded web_search_call items, came back "provider returned HTTP 400". Bisected against the gateway: any replayed item carrying an id fails with "content block web_search_tool_result is not supported on assistant messages", because the gateway's Responses-to-Claude translator rebuilds a search-and-result pair for an id it recognises and the upstream refuses the result half on an assistant message. The same item without an id is accepted and answered. Meta documents the id as optional on replay. So the encoder replays the item from the record with the sidecar's status and keeps the id in the sidecar only. PRV-3 records the rule and rejects the alternative, which is the shape Codex sends and the one that fails here. The fixture test asserts the id is absent, and the live follow-up now answers.
…with and without its id Read after the bisect rather than inferred from the error: the translator rebuilds an id-bearing item as a server_tool_use and web_search_tool_result pair on the assistant message, which the upstream refuses, and drops an id-less item entirely before Meta sees it. The second half is what makes replaying without the id harmless on this route, and what a native route would do differently.
Stage 33 slice 4. A server tool call now reaches the workspace as one event in its terminal state and one row in the tool row's grammar: `[+] web_search · <what the route reported>`, disclosure and copy of the same facts, counted among the tools. The marker wears a new seventeenth role spent on the palette's purple slot, so who ran the call is a colour and what it was is the name; a call the provider gave up wears failure like any other, which needed the outcome typed into the record. Ending states the decoder did not know (failed, searching) are now read rather than failing deserialization. The user chose the grammar and the shade from rendered candidates and reviewed the real frames at three widths on 2026-09-21; the purple and magenta slots take the values they picked. ui-ux, ENT-2, PRV-5 and the plan say what the code now does.
A live gateway finishes every web_search_call after the last message, so a row placed at `done` landed after the answer it informed, and a reopened conversation, which orders by position, showed the rows somewhere else. The Responses decoder now reports the call when its item is added; the step places it as a running row at that position and finishes it at the next revision, or reports it failed if the step ends first, keeping no block for a call that never finished. The journal holds only the finished call, so replay shows the same row at revision zero, in the same place. The activity line names the running search like a local tool's. A loopback smoke drives the burst shape the gateway produced through the binary and reopens the conversation to prove the order holds. ENT-2, PRV-5, ui-ux and the plan say what the code now does; the spike records the measurement that forced the change.
…ateway does The loopback burst now carries the reasoning items a live route sends before each message, so the smoke exercises replay beside the placed and finished server tool rows, which is the shape a captured gateway stream renders correctly once the terminal is actually read.
…the gateway or the app
…ark the stage complete
…he environment still overriding Gates run by hand before this commit (fmt, file-length, crate-graph, citations, frames, typos: all clean); the hook's mktemp is unavailable in this shell.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
A model may declare the hosted tools its route accepts (
server_tools = ["web_search"], PRV-6). The harness forwards the declaration in each dialect's spelling, decodes what the provider returns, and shows a provider-run search as a transcript row in its own colour:[~] web_search · runningwhere the provider began it, finishing once as[+] web_search · <query>or[!] … · failed(PRV-5, ENT-2). The harness never runs the search. Reopening a session shows the same rows (ENT-3).Why this shape
Decided in
.agents/spikes/provider-side-tools/against Codex, pi, grok-build and claude-code, and measured on one live gateway route. Two findings that shaped the code: the route delivers everydoneafter the last message, so the row is placed atadded; and the route delivers a whole turn in one burst after ~30 s, which is upstream of both the gateway and this app (recorded in the spike, not worked around).Verified
-D warnings, fmt, the seven pre-commit gates,scripts/smoke-server-tool.pyat three widths plus reopen, offline replay of the captured gateway stream through the real binary.8cdd911: Rust and terminal tests, static and script checks, and macOS Apple Silicon all pass.Also on this branch
feat(provider): let a route keep its bearer token in the file(8cdd911) — an inlineapi_keyon a route, with a set environment variable still overriding it, written by another agent at the owner's request. Committed on its own after the pre-commit gates were run by hand; CI covers it with the rest.