Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
9374cb4
chore(codex): upgrade openai-codex to 0.159.2
tiantt Sep 30, 2026
59c616e
fix(codex): keep replayed tool history in the order it happened
tiantt Sep 30, 2026
174feea
test(runtime): add a runtime conformance suite
tiantt Sep 30, 2026
33657ce
test(codex): add an opt-in real-model probe
tiantt Sep 30, 2026
dabd0a6
feat(codex): add model routing for direct Responses and shim transports
tiantt Sep 30, 2026
68774c5
feat(codex): add a per-turn MCP bridge for ADK tools
tiantt Sep 30, 2026
ef80ff6
test(codex): add a fake Codex for the direct provider + MCP bridge mode
tiantt Sep 30, 2026
bec12b5
feat(codex): call Responses backends directly and serve ADK tools ove…
tiantt Sep 30, 2026
3d7f17f
feat(codex): persist Codex thread rollouts in a versioned thread store
tiantt Sep 30, 2026
850c881
feat(codex): add turn-control primitives for persistent threads
tiantt Sep 30, 2026
d4b50f8
feat(codex): resume one Codex thread per session across invocations
tiantt Sep 30, 2026
5d5faf1
feat(codex): bound, steer and compact Codex turns
tiantt Sep 30, 2026
11d0144
feat(codex): end-to-end session example; treat a closed turn as cance…
tiantt Sep 30, 2026
729e0b4
fix(codex): address review findings on the direct transport
tiantt Sep 30, 2026
c6ee44f
test(codex): cover corrupt threads, bridge rejections, steer routing,…
tiantt Sep 30, 2026
7ece612
fix(codex): hand lost turns back to a resumed thread; one source of pins
tiantt Sep 30, 2026
e8099a3
fix(codex): keep header secrets off disk, retry resumes, meter the ru…
tiantt Sep 30, 2026
8f11613
feat(codex): version the thread store schema; test it on MySQL/Postgr…
tiantt Sep 30, 2026
0fcec49
fix(codex): enforce tool admission and preserve resumed execution con…
tiantt Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
90 changes: 84 additions & 6 deletions docs/content/docs/framework/agent/runtime.en.mdx

Large diffs are not rendered by default.

90 changes: 84 additions & 6 deletions docs/content/docs/framework/agent/runtime.mdx

Large diffs are not rendered by default.

12 changes: 9 additions & 3 deletions examples/codex_runtime_on_agentkit/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,9 +40,15 @@ codex_runtime_on_agentkit/
- **`openai-codex` is not a veadk dependency**, so `requirements.txt` lists it
explicitly. It pulls in `openai-codex-cli-bin`, which ships the Codex CLI
binary as a **manylinux wheel** — no separate binary install in the Linux
build. These pins mirror veadk-python's `[codex]` extra; the extra is not
used directly because uv only accepts a pre-release when its exact version is
pinned at the top level, not transitively through an extra.
build. The SDK pins that binary to its own version, so only `openai-codex`
is listed. The pin mirrors veadk-python's `[codex]` extra and is kept
explicit so the image gets this SDK even with a veadk-python release whose
extra still pins an older one.
- `veadk-python>=1.1.15` is required, not just any release with the codex
runtime: earlier releases do not set the Ark options Codex CLI 0.159 needs
(`model_reasoning_summary="none"`, `unbounded_connection_retries=false`), so
they install next to the pinned CLI and then fail every Ark call. The
`openai-codex` pin and the `veadk-python` lower bound move together.
- `fastapi` and `uvicorn` are listed too: `app.py` imports `uvicorn` directly
and the runtime's Responses→chat shim imports both at module level. They
resolve through google-adk today, but adk has been moving web deps behind
Expand Down
11 changes: 7 additions & 4 deletions examples/codex_runtime_on_agentkit/README.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,10 +36,13 @@ codex_runtime_on_agentkit/
`MODEL_AGENT_*` chat 端点(火山引擎 Ark)桥接过去。普通 Ark chat 模型无需改动即可用。
- **`openai-codex` 不是 veadk 的依赖**,所以在 `requirements.txt` 里显式列出。它会带上
`openai-codex-cli-bin`——以 **manylinux wheel** 形式打包了 Codex 二进制,Linux
构建里无需单独装二进制。它当前是 pre-release,连同其二进制依赖都**钉死到精确的
预发布版本**,这样 `uv pip install` 无需全局 `--prerelease=allow` 也能装上。
这些 pin 与 veadk-python 的 `[codex]` extra 保持一致;这里不直接用该 extra,
是因为 uv 只在**顶层**钉死精确预发布版本时才放行,通过 extra 传递则不行。
构建里无需单独装二进制。SDK 会把该二进制钉到与自身相同的版本,所以只需列出
`openai-codex`。这个 pin 与 veadk-python 的 `[codex]` extra 保持一致;显式列出是为了
在 veadk-python 已发布版本的 extra 仍钉着旧版本时,镜像也能装上这个 SDK 版本。
- `veadk-python` 要求 `>=1.1.15`,而不只是包含 codex 运行时的任意版本:更早的版本不会
设置 Codex CLI 0.159 在 Ark 上所需的选项(`model_reasoning_summary="none"`、
`unbounded_connection_retries=false`),能和钉住的 CLI 一起装上,但每次 Ark 调用都会
失败。`openai-codex` 的 pin 与 `veadk-python` 的下限必须同步调整。
- `fastapi` / `uvicorn` 也显式列出:`app.py` 直接 import `uvicorn`,runtime 的
Responses→chat shim 两者都在模块级 import。目前它们能从 google-adk 传递解析到,
但 adk 已经在把 web 依赖挪进 extra,所以 `[codex]` 和本文件都显式声明。
Expand Down
22 changes: 12 additions & 10 deletions examples/codex_runtime_on_agentkit/requirements.txt
Original file line number Diff line number Diff line change
@@ -1,16 +1,18 @@
# Installed in the image by AgentKit's default `uv pip install -r requirements.txt`.
#
# veadk-python >= 0.5.39 ships the codex runtime (veadk/runtime/codex).
# veadk-python >= 1.1.15 is the first release whose codex runtime handles
# Codex CLI 0.159 on Ark (model_reasoning_summary="none",
# unbounded_connection_retries=false). An older veadk-python paired with the
# openai-codex pin below builds a config that fails every Ark call, so the
# openai-codex pin and this lower bound must move together.
#
# This mirrors veadk-python's own `[codex]` extra. The extra is not used
# directly because uv refuses a *transitive* pre-release: openai-codex and its
# bundled-binary dependency openai-codex-cli-bin (the Codex CLI as a manylinux
# wheel) are pre-releases, and uv only accepts them when the exact pre-release
# version is pinned at the top level, as below. Keep these pins in sync with
# the `[codex]` extra in pyproject.toml.
veadk-python>=0.5.39
openai-codex==0.1.0b3
openai-codex-cli-bin==0.137.0a4
# This mirrors veadk-python's own `[codex]` extra, pinned here so the image
# gets this exact SDK even with a veadk-python release whose extra still pins
# an older one. openai-codex pins its matching Codex CLI binary
# (openai-codex-cli-bin) exactly, so the binary is not listed separately. Keep
# this pin in sync with the `[codex]` extra in pyproject.toml.
veadk-python>=1.1.15
openai-codex==0.159.2

# The Responses->chat shim (veadk/runtime/codex/proxy.py) imports these at
# module level, and app.py imports uvicorn directly. They resolve transitively
Expand Down
19 changes: 19 additions & 0 deletions examples/codex_session_lifecycle/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Codex session lifecycle

One `runtime="codex"` session, end to end: an MCP tool, a Python function
tool and a skill; sandboxed file writes; a streamed answer; `runner.steer()`
into a running turn; cancelling a running turn; and resuming the session's
Codex thread on the next request.

```bash
pip install "veadk-python[codex]"
export MODEL_AGENT_API_KEY=... MODEL_AGENT_API_BASE=https://ark.cn-beijing.volces.com/api/v3 MODEL_AGENT_NAME=...
python examples/codex_session_lifecycle/main.py
```

- Ark is called directly and the agent's tools reach Codex over a local MCP
server, so Codex drives the tool loop (`model_transport="auto"`).
- Each session keeps one Codex thread (`thread_mode="resume"`), saved with the
session in SQLite here, so a later request — or a second run of the script —
resumes it instead of replaying the transcript.
- Reuses the skill and MCP server of `examples/codex_with_skill_and_mcp`.
13 changes: 13 additions & 0 deletions examples/codex_session_lifecycle/README.zh.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Codex 会话全流程

一个 `runtime="codex"` 会话的完整流程:MCP 工具、Python 函数工具与 skill;沙箱内写文件;流式输出;用 `runner.steer()` 向进行中的回合追加指令;取消进行中的回合;下一次请求恢复该会话的 Codex thread。

```bash
pip install "veadk-python[codex]"
export MODEL_AGENT_API_KEY=... MODEL_AGENT_API_BASE=https://ark.cn-beijing.volces.com/api/v3 MODEL_AGENT_NAME=...
python examples/codex_session_lifecycle/main.py
```

- Codex 直接调用方舟,Agent 的工具经本地 MCP server 交给 Codex,由 Codex 驱动工具循环(`model_transport="auto"`)。
- 每个会话保有一个 Codex thread(`thread_mode="resume"`),此处随会话存在 SQLite 中;之后的请求或再次运行脚本都会恢复该 thread,而不是回放对话记录。
- 复用 `examples/codex_with_skill_and_mcp` 的 skill 与 MCP server。
178 changes: 178 additions & 0 deletions examples/codex_session_lifecycle/main.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,178 @@
# Copyright (c) 2025 Beijing Volcano Engine Technology Co., Ltd. and/or its affiliates.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

"""A `runtime="codex"` session end to end: tools, sandbox, streaming, steering,
cancellation and resume.

One session, three invocations:

1. Plan a trip. Codex calls an MCP tool (weather) and a Python function tool
(city research), follows a skill's reply style, and writes the plan to
`plan.md` in its sandboxed workspace. The answer streams as it is written.
While the research tool runs, `runner.steer()` adds an instruction to the
turn in flight.
2. Start a second request and cancel it: the Codex turn is interrupted.
3. Ask about the earlier work: the session's Codex thread is resumed, so Codex
answers from its own history and the files it wrote.

Sessions live in SQLite, so the Codex thread (saved with the session) also
survives a process restart: run the script twice to see turn 3 of the first
run's session resumed by the second.

Run:
python examples/codex_session_lifecycle/main.py

Requires ``pip install "veadk-python[codex]"`` and a Responses-capable model
(Volcengine Ark or OpenAI) via ``MODEL_AGENT_API_KEY`` / ``MODEL_AGENT_API_BASE``
/ ``MODEL_AGENT_NAME``.
"""

import asyncio
import os
import sys
from pathlib import Path

from google.adk.skills import load_skill_from_dir
from google.adk.tools.mcp_tool.mcp_session_manager import StdioServerParameters
from google.adk.tools.mcp_tool.mcp_toolset import MCPToolset
from google.adk.tools.skill_toolset import SkillToolset
from google.genai import types

from veadk import Agent, Runner
from veadk.memory.short_term_memory import ShortTermMemory
from veadk.runtime.codex import current_workspace

# Reuse the sibling example's skill and MCP server.
_SIBLING = Path(__file__).resolve().parent.parent / "codex_with_skill_and_mcp"
_SESSION_ID = "trip-planning"
_DATABASE = "/tmp/veadk_codex_session_lifecycle.db"

# Set while the research tool runs, so the demo can steer the turn right then.
_researching = asyncio.Event()


async def research_city(city: str) -> dict:
"""Look up the must-see places of a city.

Args:
city (str): The city to research.

Returns:
dict: Highlights of the city.
"""
_researching.set()
await asyncio.sleep(3) # a slow lookup: the window in which we steer
workspace = current_workspace() # the turn's sandbox directory, if any
return {
"city": city,
"highlights": ["Forbidden City", "Temple of Heaven", "Hutong walk"],
"workspace": workspace,
}


def build_agent() -> Agent:
return Agent(
name="trip_planner",
description="Plans short city trips.",
instruction=(
"Plan trips. Use the weather tool and research_city, then save the "
"plan as plan.md in the working directory with a shell command."
),
runtime="codex",
model_name=os.getenv("MODEL_AGENT_NAME", "deepseek-v4-pro-260425"),
model_api_base=os.getenv(
"MODEL_AGENT_API_BASE", "https://ark.cn-beijing.volces.com/api/v3"
),
model_api_key=os.getenv("MODEL_AGENT_API_KEY"),
tools=[
SkillToolset(
skills=[load_skill_from_dir(str(_SIBLING / "skills" / "weather-style"))]
),
MCPToolset(
connection_params=StdioServerParameters(
command=sys.executable, args=[str(_SIBLING / "mcp_server.py")]
)
),
research_city,
],
codex_runtime_config={
# Defaults, spelled out: Ark is called directly and the session's
# Codex thread is resumed on every turn.
"model_transport": "auto",
"thread_mode": "resume",
"sandbox": "workspace_write",
"turn_timeout_seconds": 300,
},
)


async def ask(runner: Runner, text: str) -> None:
print(f"\nUser: {text}\nAgent: ", end="", flush=True)
async for event in runner.run_async(
user_id=runner.user_id,
session_id=_SESSION_ID,
new_message=types.Content(role="user", parts=[types.Part(text=text)]),
):
for call in event.get_function_calls() or []:
print(f"\n [tool] {call.name}", flush=True)
if not event.content or not event.content.parts:
continue
for part in event.content.parts:
if part.text and not part.thought and event.partial:
print(part.text, end="", flush=True) # stream the answer
print()


async def main() -> None:
runner = Runner(
agent=build_agent(),
short_term_memory=ShortTermMemory(
backend="sqlite", local_database_path=_DATABASE
),
)
# Reuse the session across runs: its Codex thread is stored with it.
session = await runner.session_service.get_session(
app_name=runner.app_name, user_id=runner.user_id, session_id=_SESSION_ID
)
if session is None:
await runner.short_term_memory.create_session(
app_name=runner.app_name, user_id=runner.user_id, session_id=_SESSION_ID
)

# 1. Tools, skill, sandbox, streaming -- and a steer mid-turn.
turn = asyncio.create_task(
ask(runner, "Plan a 2-day Beijing trip and save it to plan.md.")
)
await _researching.wait()
delivered = await runner.steer(_SESSION_ID, "Also add a short packing list.")
print(f"\n [steer] delivered to the running turn: {delivered}")
await turn

# 2. Cancel a request whose Codex turn is running (it is inside the research
# tool): the turn is interrupted, and the next request still works.
_researching.clear()
turn = asyncio.create_task(ask(runner, "Now plan Shanghai the same way."))
await _researching.wait()
turn.cancel()
try:
await turn
except asyncio.CancelledError:
print("\n [cancel] the Shanghai request was cancelled")

# 3. Resume: Codex answers from its own thread and workspace.
await ask(runner, "What did you save earlier, and what is in plan.md?")


if __name__ == "__main__":
asyncio.run(main())
17 changes: 10 additions & 7 deletions examples/codex_with_skill_and_mcp/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,10 +46,13 @@ API — so the two tools take different paths:
- **Skill** → materialized into Codex's on-disk skill directory
(`$CODEX_HOME/skills/<name>/SKILL.md`) and discovered by Codex's native skill
system. Backend-independent.
- **MCP tool** → Codex can't be handed MCP tools directly (it presents them to
the model as a `namespace` tool the chat backend rejects), so the runtime's
Responses shim advertises them to the backend as plain `function` tools and
executes them itself, invisibly to Codex.
- **MCP tool** → handed to Codex as one of its own tools. On a
Responses-capable backend (Volcengine Ark, OpenAI) Codex calls the model
directly and reaches the agent's tools through a local MCP server the runtime
runs for the turn, so Codex drives the tool loop. For a chat-only backend
the runtime's Responses shim sits in between and executes the tools itself.
`CodexRuntimeConfig(model_transport=...)` (`auto` / `direct` / `shim`)
chooses; `auto` picks direct for Ark and OpenAI.

Both are handled by the runtime — the agent code is just normal tool wiring.

Expand All @@ -67,9 +70,9 @@ python examples/codex_with_skill_and_mcp/main.py

## Notes

- Tools are dispatched by the runtime shim, while calls, results, state
changes, confirmations, and authentication surface as standard ADK events
for Session/Trace/UI.
- Tools execute in the runtime (through the MCP bridge or the shim), and
calls, results, state changes, confirmations, and authentication surface as
standard ADK events for Session/Trace/UI.
- Static authentication (headers / bearer tokens / ve-identity workload
tokens) and ADK interactive authentication requested during tool execution
are supported. Authentication required before an MCP toolset can list tools
Expand Down
4 changes: 2 additions & 2 deletions examples/codex_with_skill_and_mcp/README.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ Beijing: sunny, 28°C. Have a nice day!
Codex 接管了整轮(而不是 ADK 的 LLM flow),且只会说 Responses API——所以两个工具走不同的路:

- **Skill** → 被物化到 Codex 的磁盘 skill 目录(`$CODEX_HOME/skills/<name>/SKILL.md`),由 Codex 原生 skill 机制发现。与后端无关。
- **MCP 工具** → 不能直接交给 Codex(它会把 MCP 工具以 `namespace` 类型呈现给模型,而 chat 后端不认),所以由 runtime 的 Responses shim 把它们当普通 `function` 工具喂给后端、并**自己执行**,对 Codex 不可见。
- **MCP 工具** → 作为 Codex 自己的工具交给它。后端支持 Responses API(火山方舟、OpenAI)时,Codex 直接调用模型,并通过 runtime 为本轮启动的本地 MCP server 调用 agent 的工具,由 Codex 驱动工具循环;后端只支持 chat 时,由 runtime 的 Responses shim 居中并**自己执行**工具。用 `CodexRuntimeConfig(model_transport=...)`(`auto` / `direct` / `shim`)选择,`auto` 对方舟和 OpenAI 选直连。

这些都由 runtime 处理——Agent 代码就是普通的工具挂载。

Expand All @@ -57,7 +57,7 @@ python examples/codex_with_skill_and_mcp/main.py

## 说明

- 工具由 runtime 的 shim 调度,但调用、结果、状态变更、确认和鉴权都会作为标准 ADK 事件进入 Session/Trace/UI。
- 工具在 runtime 中执行(经 MCP bridge 或 shim),调用、结果、状态变更、确认和鉴权都会作为标准 ADK 事件进入 Session/Trace/UI。
- 支持静态鉴权(header / bearer token / ve-identity workload token)以及工具执行中触发的 ADK 交互式鉴权;MCP toolset 在列举工具前触发的鉴权仍取决于对应 ADK/MCP 客户端能力。
- `runtime="codex"` 是**沙箱执行运行时**,不是 ADK 执行流程的等价替代品。`Agent` 上有一部分配置在它下面会**直接报错**(`sub_agents`、`output_schema`、`planner`、`code_executor`、`system_instruction` 以外的 `generate_content_config`、`include_contents="none"`、`enable_supervisor`,以及显式传入的 `model=`),另一部分会被丢弃并告警(`knowledgebase`、`example_store`、`skills_mode` 等)。详见[支持矩阵](../../docs/content/docs/framework/agent/runtime.mdx#支持矩阵)。
- 注意本例依赖的区别:ADK 的 `SkillToolset` 会被桥接进 Codex 原生 skill 系统,但 VeADK 自己的 `Agent(skills_mode=...)` **不会**——后者只会告警且不生效。
16 changes: 9 additions & 7 deletions examples/codex_with_skill_and_mcp/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,14 +14,15 @@

"""A `runtime="codex"` agent that uses both a local skill and an MCP tool.

On a Codex runtime backed by a chat model (e.g. Volcengine Ark):
On a Codex runtime:

- **Skills** are materialized into Codex's on-disk skill directory and driven by
Codex's native skill system.
- **MCP / function tools** can't be handed to Codex directly (Codex presents
them to the model as a `namespace` tool the chat backend rejects), so the
runtime's shim advertises them to the backend as plain functions and executes
them itself.
- **MCP / function tools** reach Codex as its own tools. On a Responses-capable
backend (Volcengine Ark, OpenAI) Codex calls the model directly and gets the
agent's tools through a local MCP server the runtime runs for the turn, so
Codex drives the tool loop. For a chat-only backend the runtime's shim sits
in between and executes the tools itself.

Both are just normal VeADK/ADK wiring — the runtime handles the rest.

Expand Down Expand Up @@ -60,7 +61,7 @@ def build_agent() -> Agent:
skill_toolset = SkillToolset(skills=[load_skill_from_dir(str(_SKILL_DIR))])

# MCP: a stdio MCP server launched as a subprocess. The codex runtime lists
# its tools and executes them via the shim. Swap StdioServerParameters for
# its tools and hands them to Codex (see above). Swap StdioServerParameters for
# StreamableHTTPConnectionParams(url=...) to point at a remote MCP server.
weather_mcp = MCPToolset(
connection_params=StdioServerParameters(
Expand Down Expand Up @@ -96,7 +97,8 @@ async def main() -> None:
session_id="s1",
new_message=types.Content(role="user", parts=[types.Part(text=question)]),
):
if not event.content or not event.content.parts:
# Partial events are streaming chunks of text the final event repeats.
if event.partial or not event.content or not event.content.parts:
continue
for part in event.content.parts:
if part.text and not part.thought:
Expand Down
Loading
Loading