Skip to content

feat(codex): run Codex on the official SDK with direct Responses, MCP tools and resumable threads - #1153

Open
tiantt wants to merge 19 commits into
mainfrom
chore/upgrade-openai-codex-sdk
Open

tiantt wants to merge 19 commits into
mainfrom
chore/upgrade-openai-codex-sdk

Conversation

@tiantt

@tiantt tiantt commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Move Agent(runtime="codex") to the official openai-codex==0.159.2 SDK, with direct Responses transport, MCP access to ADK tools, and resumable Codex threads. Codex owns the agent loop and sandbox; VeADK retains session binding, tool access, event translation, tracing, and execution controls.

Changes

  • Select direct Responses transport for compatible Ark, BytePlus and OpenAI endpoints; retain the shim for chat-only backends and configurations that require request-body passthrough.
  • Persist versioned thread rollouts with the session, with resume retries, size limits and bounded in-memory storage. Resume backfills missing messages and completed tool calls/results after a lost save, using ADK's current-branch filtering.
  • Cancel and drain active shim requests before releasing tool resources, including requests running on another event loop. Revoke requests on service shutdown as well.
  • Enforce the ADK/MCP tool budget before execution. Both transports count individual calls; direct rejects excess calls and the shim rejects a parallel batch when it cannot admit the whole batch. Exhaustion is surfaced as CodexToolIterationLimitError.
  • Apply before-model callback model selection to Codex configuration and backend requests. Keep credentials out of config files and sandboxed shell environments.
  • Provide turn timeouts, interruption, steering, auto compaction, event translation and OpenTelemetry metrics. Remove the differential test suite's ambiguous conftest import.

Compatibility

The public Agent(runtime="codex") entry point is unchanged. Direct transport and resumable threads are selected by default when supported. Set model_transport="shim" or thread_mode="ephemeral" to opt out. The default turn timeout is 1800 seconds.

max_tool_iterations now counts individual bridged ADK/MCP calls on both transports, including each call in a parallel batch; native Codex tools are excluded. Workspace files remain instance-local, and cross-instance execution leases are not implemented.

Validation

  • Added 11 regression cases for cancellation and draining, parallel tool admission, lost-save recovery, branch isolation, and callback-selected models.
  • Submission checks: pre-commit (ruff check, ruff format, gitleaks) and git diff --cached --check passed.
  • Relevant unit tests: 532 passed, 38 skipped, 2 xfailed.
uv run --no-sync pytest \
  tests/runtime/codex tests/runtime/differential \
  tests/runtime/test_model_callbacks.py \
  tests/cli/test_codex_app_server.py \
  tests/cli/test_codex_presentation_events.py -q -rs -n 4

Real Codex binary smoke tests, the paid model probe, and external MySQL/PostgreSQL tests were not run for this update; their opt-in cases account for the skips. GitHub CI is evaluated separately.

tiantt and others added 15 commits September 30, 2026 13:30
The SDK is stable now and pins its matching Codex CLI binary exactly,
so only `openai-codex` is pinned (in the extra and in the AgentKit
example requirements).

- Pass `personality` / `effort` as plain strings instead of importing
  the SDK's private `openai_codex.generated` enums.
- Pin three settings in the generated config.toml that CLI 0.159
  changed underneath the runtime:
  - `unbounded_connection_retries = false`: an unreachable shim was
    retried forever, hanging the invocation (no turn timeout exists).
  - `goals = false`: goal tools were advertised to the backend model
    although every thread here is ephemeral.
  - `model_reasoning_summary = "none"`: Ark's Responses API rejected
    every request over `reasoning.summary` ("unknown field").
- Explicitly ignore the four notification types the new SDK adds; none
  of them can leave a turn waiting on the runtime.
- Add baseline tests for cancellation (while waiting on the model,
  during a tool, with a failing interrupt), the pinned config settings,
  the SDK/binary version pairing, and a real-binary smoke test that an
  unreachable shim fails the turn in bounded time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two ways the history the model sees drifted from what actually happened,
both older than the SDK upgrade (reproduced on 0.1.0b3 as well). On a
multi-step investigation they made the model repeat itself until the
call budget ran out.

- The shim appended its ADK tool pairs to the tail of every request, so
  the ADK results always read as the model's most recent step. It kept
  announcing it had just fetched its data and re-ran the same analysis.
  Each pair is now anchored to the Codex-visible function call of the
  reply it preceded and spliced back in before that reply; pairs whose
  anchor is gone (e.g. after compaction) still fall back to the tail.
- Commands were recorded with Codex's login-shell wrapper
  (`/bin/zsh -lc '...'`). The next invocation's model saw that in its
  replayed history, copied it into its own commands, and Codex wrapped
  them again. The recorded command is now the one the model sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pins the observable contract every agent runtime must keep, so the Codex
runtime redesign (thread resume, MCP tool bridge, steer) can change its
internals against a fixed bar. Scenarios run offline through the real
Agent and Runner for adk, codex and piagent via per-runtime adapters;
a new runtime only adds an adapter and its capability set.

- Required: single turn, streaming, multi-turn, session isolation,
  cancellation (waiting on the model and mid-tool), tool call, tool
  failure, concurrent sessions, usage, lifecycle cleanup.
- Capability-gated: approvals, MCP tools, skills; and resume across a
  restart, steer, turn timeout and compaction, written against adapter
  hooks no runtime implements yet, so skipped until a runtime declares
  them.
- Two real gaps are strict xfails with their cause: runtime="adk" lets a
  raising tool crash the invocation, and runtime="piagent" never finishes
  a turn cancelled while a bridged tool runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fixed two-turn scenario (one ADK tool, three shell steps, then a
follow-up on the same workspace) that a Codex runtime PR can run once
against a real model for ~75K tokens, instead of full examples that cost
millions. It fails on the regressions seen so far: backend-rejected
request fields, looping on out-of-order tool history, the model copying
Codex's shell wrapper, tools re-run, lost multi-turn context, and a
token total over CODEX_PROBE_MAX_TOKENS.

Skipped unless CODEX_RUN_PROBE=1 and MODEL_AGENT_* are set; never in CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex speaks the Responses API, and Ark / BytePlus ModelArk / OpenAI serve
it compatibly, so for those hosts Codex can call the model directly instead
of going through the in-process Responses-to-chat shim. This adds the
routing layer the runtime integration will use: resolve_transport picks
"direct" or "shim" (explicit CodexRuntimeConfig.model_transport, env
VEADK_CODEX_MODEL_TRANSPORT, or host-based auto), and direct_route /
shim_route build the provider as a thread_start(config=...) override.

The key only travels in route.env (excluded from repr), never in the
provider config Codex is handed. lean_codex_config pins the settings the
shim path needed (no reasoning summary, bounded connection retries) and
trims tools a VeADK turn cannot use (multi_agent, web_search, view_image,
goals, request_user_input), roughly halving input tokens per request.
Every key was checked against CLI 0.159.2 at thread level by capturing
the request the real binary sends; an opt-in codex_smoke test keeps that
proof.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ADK tools currently run inside the Responses shim, invisible to Codex, so
the shim has to own the tool loop and its history. Serving them to Codex
as an MCP server injected per thread lets Codex own the loop instead. This
adds that server as a standalone module; wiring it into the runtime is a
separate change.

One in-process streamable-HTTP server per event loop (get_bridge) serves
every turn. Each turn registers its specs and executors and gets a random
bearer token; tools/list and tools/call only ever see that token's tools,
and unknown tokens get 401. tools/call runs the executor under the turn's
OTel context with the model's call_id from `_meta.callId`, returns the
executor's JSON as structuredContent, records interrupt statuses
(pending / authentication_required / confirmation_required / transferred)
for the runtime to stop the turn, and turns executor exceptions into
isError results. A dropped HTTP request or `notifications/cancelled`
cancels the executor.

The transport answers with plain JSON, not SSE: sse-starlette latches a
process-global exit flag when it sees a uvicorn server shut down with a
stream open, after which every later SSE response ends immediately, so
one bridge stop mid-call would break every later bridge. The server also
no longer captures the host's SIGINT/SIGTERM (uvicorn's capture_signals).

codex_server_config() emits the Codex keys verified against 0.159.2:
default_tools_approval_mode="approve" (deny_all blocks MCP calls
otherwise), supports_parallel_tool_calls, tool_timeout_sec (above the ADK
timeout), startup_timeout_sec and required=true.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The codex runtime is moving to a direct mode: Codex calls the model
provider itself (thread config model_providers.<id>) and reaches ADK tools
through VeADK's local streamable-HTTP MCP bridge (mcp_servers.veadk).
ShimDrivingCodex can only exercise the Responses shim, so none of that
path was testable offline.

DirectDrivingCodex mimics what codex 0.159.2 was observed to do in the
MCP spikes: it reads the provider and MCP servers from
thread_start(config=...) and credentials from CodexConfig.env, calls the
model through the patched litellm.aresponses with no shim, connects to each
server with the real mcp streamable-HTTP client (bearer from the env var),
advertises the tools as a {"type": "namespace", "name": "mcp__<server>"}
tool, executes namespaced function calls via tools/call with _meta.callId
(concurrently when supports_parallel_tool_calls), feeds results back with
codex's "Wall time ... Output:" framing, and emits mcpToolCall item
notifications, per-model-call token usage and turn/completed (interrupted
when interrupt() cancels in-flight model/MCP calls). Notification payloads
carry every field the real SDK models require, so real models are built
when openai-codex is installed.

ScriptedBackend gains, additively, the ability to answer with a
namespaced call when a plan names a tool only advertised inside a
namespace tool, and records namespaced tool names as plain names so the
two arms stay comparable.

The self-tests run the fake against a real FastMCP server on loopback and
reset sse_starlette's process-wide AppStatus.should_exit, which otherwise
latches after the first uvicorn shutdown and hangs every later bridge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…r MCP

Until now every Codex turn went through the in-process Responses->chat
shim, which also ran the agent's ADK tools in a loop Codex could not
see. For backends that speak the Responses API (Ark, BytePlus, OpenAI)
neither is needed: a spike showed Codex calls Ark directly with none of
the shim's compatibility patches, and reaches ADK tools through a local
MCP server. That makes Codex the single owner of the tool loop, which
the planned thread resume depends on (shim-run tools never reach the
Codex rollout).

- `model_transport` (auto/direct/shim, env VEADK_CODEX_MODEL_TRANSPORT)
  picks the path; `auto` is direct for Ark, BytePlus and OpenAI hosts
  and keeps the shim for everything else.
- Direct: the provider and a trimmed tool set go in the thread config,
  the key only in the subprocess env; ADK tools are registered on the
  MCP bridge under a per-turn bearer token.
- Codex's mirror of a bridged MCP call is dropped, so each ADK tool call
  appears once, with callbacks and state deltas applied.
- `max_llm_calls` is charged per usage update (Codex announces nothing
  before a request), and the turn is interrupted once it is exceeded.
- A bridged tool that needs a credential or confirmation, or transfers
  control, interrupts the turn instead of letting the model continue.
- The conformance suite gains a codex-direct adapter; the example and
  runtime docs describe both transports.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
runtime="codex" is moving to one persistent Codex thread per VeADK session
(thread_start(ephemeral=False), then thread_resume). In production (AgentKit)
turns land on different stateless instances with ephemeral disks, and a
thread's whole context lives in one rollout JSONL file under CODEX_HOME;
copying that file into a fresh CODEX_HOME is enough to resume. So VeADK has
to carry rollouts between instances itself.

- rollout_io: find/export/import a thread's rollout. Import is atomic
  (tmp + fsync + rename), 0600, and refuses relpaths outside
  CODEX_HOME/sessions or not named for the thread, so stored data cannot
  overwrite config.toml/auth.json or escape via symlinked dirs.
- thread_store: CodexThreadStore keyed by (app, user, session, agent) with
  compare-and-set versions so racing instances cannot silently drop a turn,
  plus the instruction hash the thread started with. Backends: in-memory,
  local dir (flock), and a SQL table (gzip blob, LONGBLOB on MySQL) on the
  short-term memory's own DatabaseSessionService.db_engine, so sqlite,
  mysql and postgresql (including its search_path schema) are covered with
  no extra configuration. select_thread_store picks it from ShortTermMemory.

Nothing is wired into runtime.py yet. Rollout contents are never logged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Keeping one Codex thread per session makes the app-server's turn semantics
load-bearing, and against Codex 0.159.2 they are hostile to naive use:
turn() on a thread with an active turn joins it (new input lost when it
races an interrupt), interrupt() right after turn() is rejected with -32600
"no active turn" and the turn then runs to completion, the interrupted
turn stays active until its turn/completed arrives 10-70 ms later, and
compact() only starts a compaction turn whose events reach no handle and
which rejects turn() with -32603 ActiveTurnNotSteerable.

turn_control.py turns that into small primitives the runtime can wire in
without further protocol knowledge: a per-session asyncio lock that drops
idle entries, an active-turn registry whose steer() never starts a turn,
TurnCompletion fed by the stream pump, interrupt_turn() that retries the
early rejection and waits for turn/completed, start_fresh_turn() that
refuses a joined turn and waits out a compaction, compact_and_wait() that
polls thread.read(include_turns=True), and run_with_turn_timeout() raising
CodexTurnTimeout. Errors are classified through the SDK's public
InvalidRequestError / InternalRpcError. Not wired into runtime.py yet.

Unit tests cover every branch on fakes; codex_smoke tests (CODEX_RUN_SMOKE=1)
check the five behaviours against the real binary with a stub backend.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every invocation used to start an ephemeral Codex thread and replay the
session as a JSON transcript in the prompt. Codex then never saw its own
earlier commands, outputs and plans, and the replayed commands taught
the model to copy Codex's shell wrapper. With ADK tools now reaching
Codex over MCP (so they are in its history), the thread itself can be
kept.

- `thread_mode` (resume/ephemeral, env VEADK_CODEX_THREAD_MODE), resume
  by default on the direct transport; the shim stays ephemeral.
- The thread's rollout is saved to the thread store after every turn
  (failures and cancellation included) and imported into the next
  invocation's fresh CODEX_HOME, so any instance can resume it.
- A resumed turn sends the current message only, plus what the user or
  other agents said since this agent last replied, and the results of
  tool calls that ran once the user confirmed them.
- Settings are passed again on every resume: Codex falls back to its
  defaults for model and sandbox otherwise.
- A changed instruction starts a new thread from the transcript (Codex
  ignores new developer instructions on resume); a failed resume or a
  corrupt record does the same instead of failing the turn.
- A per-session lock keeps two invocations off one thread; a lost
  optimistic save is logged, not raised.
- Usage is reported per turn, not as the resumed thread's running total.
- The offline fake persists and resumes threads; the conformance suite's
  codex-direct adapter now declares resume_across_restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With persistent threads a turn has to end cleanly and on time, can be
redirected while it runs, and needs its history kept in check.

- `turn_timeout_seconds`: past the deadline the turn is interrupted,
  given a grace period, and the invocation fails with a TimeoutError.
  The watchdog owns the pump, so a pump it has to cancel surfaces as the
  timeout rather than as a cancellation.
- Cancellation waits until the Codex turn has really stopped (retrying
  an interrupt that lands before the model request starts), so the next
  invocation never joins a dying turn.
- `Runner.steer(session_id, text)` adds an instruction to the turn in
  flight through a generic `BaseRuntime.steer`; the Codex runtime
  registers its running turn per session (direct transport only: a
  steered message through the shim would carry no turn marker).
- `auto_compact_token_limit` hands compaction to Codex itself
  (`model_auto_compact_token_limit`); no summarizer of our own.
- The turn's span records `veadk.codex.*` metadata: transport, thread
  and turn ids, resume, model, sandbox, approval mode, tool count,
  status and duration -- no prompts, arguments or keys.
- The fake app-server steers and compacts like Codex 0.159; the
  conformance suite's codex adapters now declare turn_timeout, and
  codex-direct also steer and compaction. Docs gain an architecture
  section with the local/managed backend roadmap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lled

- examples/codex_session_lifecycle: one session through MCP and function
  tools, a skill, sandboxed writes, streaming, a steer, cancelling a
  running turn, and resuming the session's Codex thread. Verified
  against Ark.
- When the caller's task is cancelled, ADK's Runner closes the runtime
  generator (GeneratorExit) instead of raising CancelledError into it.
  That path logged the turn as failed and never interrupted Codex; it
  is now a cancellation with the same safe interrupt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From an anr code review of the Codex runtime work:

- Credentials leaked into the sandboxed shell: `env` run by Codex printed
  the direct transport's model key (and would show the shim turn token
  and the MCP bridge token). The generated config now pins a shell
  environment policy excluding VEADK_CODEX_*, *API_KEY*, *SECRET* and
  *TOKEN*; a real-binary smoke test runs `env` and checks the key is gone.
- `max_tool_iterations` only bounded the shim: on the direct transport
  the runtime now counts the MCP bridge's calls and interrupts the turn
  with CodexToolIterationLimitError past the limit.
- `extra_body` was silently dropped on the direct transport. Under
  `auto` an agent with a body to forward stays on the shim; an explicit
  `direct` logs a warning. VeADK's default Ark body (caching) is dropped
  on Responses anyway and does not force the shim.
- `turn_timeout_seconds` defaults to 1800: an unbounded turn would also
  block its session's queued invocations.
- Saving the rollout no longer swallows cancellation, and cleanup always
  runs; rollout import/export move off the event loop.
- The backfill handed to a resumed thread is capped (message count,
  message size, resumed tool result size).
- Failure logs carry thread/turn/session ids and a bounded error message;
  a codex_turn_started line correlates turn ids with invocations.
- Docs and the AgentKit example (veadk-python>=1.1.15, moved together
  with the openai-codex pin) describe the above.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… save races

A review of the runtime="codex" work found guarded behaviour with no test
pinning it, so a regression in any of these would have gone unnoticed:

- Thread stores: LocalDir and Database now must raise ThreadStoreCorrupt
  for an undecodable blob or a sha mismatch, and _load_thread must delete
  the corrupt record so the next expected_version=None save succeeds
  (otherwise the session is pinned to throwaway threads for good).
- MCP bridge: oversized bodies get 413 before any executor runs, paths
  other than /mcp get 404, and non-loopback Host/Origin are refused
  (DNS-rebinding protection), each paired with a loopback positive.
- Config: VEADK_CODEX_THREAD_MODE accepted/rejected values, invalid
  thread_mode, and zero/negative turn_timeout_seconds and
  auto_compact_token_limit.
- Runner.steer: routing to a nested codex sub-agent with the runner's
  app_name and default user_id, False when nothing delivers, first
  delivery wins, and a pure adk tree never consults a runtime.
- Cross-instance saves: a ThreadStoreConflict on save is logged as
  codex_thread_save_conflict and does not fail the invocation, and a turn
  after another writer's save resumes that writer's record and version.
- Conformance compaction: the compaction request advertises no agent
  tools while the agent's own turns do, and the next turn no longer
  carries turn 1's answer. The module docstring now lists which adapters
  declare steer/turn_timeout/compaction/resume_across_restart.

Every new assertion was mutation-checked against a temporarily broken
veadk/ (or, for the compaction tool list, the fake Codex) and restored
byte-identically.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@tiantt tiantt changed the title chore(codex): upgrade openai-codex to 0.159.2 and fix tool-history replay order feat(codex): run Codex on the official SDK with direct Responses, MCP tools and resumable threads Sep 30, 2026
tiantt and others added 4 commits September 30, 2026 18:26
From the semantic axes of an anr review:

- A turn whose rollout save was lost (cross-instance conflict, store
  error, cancellation) vanished from the Codex thread: the backfill
  anchored on this agent's last reply, and that reply belonged to the
  lost turn. Records now store the last invocation their rollout covers
  (`covered_invocation_id`), and a resumed turn is handed everything
  after it, including this agent's own lost replies.
- A codex agent transferred to within the same invocation lost what the
  parent agent said in it: the backfill dropped the whole current
  invocation. It now skips only the current user message (rendered as
  the prompt) and this agent's own events.
- The settings pinned for Codex lived twice (hand-written config.toml
  and the direct thread override) and had already drifted.
  `pinned_codex_settings()` is now the single source for both.
- Docs: the per-session lock applies to direct + resume only;
  `max_tool_iterations` counts individual calls on the direct transport.
  turn_control's module doc now describes the real wiring and names the
  helpers that are not wired in.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ntime

From a second round of review on the Codex runtime:

- Sensitive model headers (Authorization, Cookie, or names with key /
  token / secret / password / signature) no longer land in Codex's
  config file on the direct transport: they go through
  `env_http_headers`, with values only in VEADK_CODEX_* env vars the
  sandboxed shell cannot see. A real-binary test checks the header
  reaches the backend and appears in neither the file nor the shell.
- Resume no longer abandons a thread's native context on any error:
  transient overload is retried (bounded, with backoff); unknown threads
  and invalid parameters still fall back to a new thread at once.
- A resumed turn on an empty workspace (new instance, restart) is told
  its earlier files may be gone.
- Rollouts over `VEADK_CODEX_MAX_ROLLOUT_BYTES` (32 MiB) are refused;
  the binding is dropped so the next turn starts fresh. The in-process
  store keeps rollouts compressed and evicts least-recently-used ones.
- OpenTelemetry metrics with low-cardinality attributes: resume and save
  outcomes, turns by status/transport with duration, startup time, and
  tokens.
- Docs: limitations (workspace not persisted, no cross-instance lease,
  best-effort max_llm_calls on direct, rollout cap), sensitive headers,
  resume retries, metrics and the new env vars.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eSQL

Every stored thread record now carries a schema version (a column on
veadk_codex_threads, a field in the local-dir header). A record written by a
newer VeADK is neither read nor overwritten: load raises
ThreadStoreIncompatible, the update CAS excludes newer rows, and the runtime
runs the turn on a new thread without saving (resume outcome
"incompatible"). A table created by a pre-release build without the column
fails loudly with ThreadStoreSchemaError.

The store contract tests and the real-binary resume test now also run
against a real MySQL / PostgreSQL when VEADK_TEST_MYSQL_URL /
VEADK_TEST_POSTGRES_URL are set, each test on its own table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…text

Co-authored-by: TRAE CLI <traecli@bytedance.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant