diff --git a/LICENSE b/LICENSE
new file mode 100644
index 0000000..7c254dc
--- /dev/null
+++ b/LICENSE
@@ -0,0 +1,21 @@
+MIT License
+
+Copyright (c) 2026 xBot contributors
+
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to deal
+in the Software without restriction, including without limitation the rights
+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+
+The above copyright notice and this permission notice shall be included in all
+copies or substantial portions of the Software.
+
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+SOFTWARE.
diff --git a/README.md b/README.md
index dbbdaf2..d154fcd 100644
--- a/README.md
+++ b/README.md
@@ -2,16 +2,22 @@
-
-**Your own AI coworkers, on your own Mac.**
+
Your own AI coworkers, on your own Mac
Create agents, give them a computer, watch them work, and take the wheel when you want to.
-Bring any model — OpenAI, Anthropic, Google, xAI, or a model running locally through Ollama.
-Your agents, their files, and their browsers stay on this Mac.
+Bring any model: OpenAI, Anthropic, Google, xAI, or a model running locally through Ollama.
+Your agents, their files, their browsers and your conversations stay on this Mac.
+
+*Not released yet: there is no signed download. Build it from source below, and see the
+[launch checklist](docs/13-launch-checklist.md) for what stands between here and a `.dmg`.*
-*In v1 your conversation history is the exception: it is stored by CopilotKit, the service xBot's
-engine is built on. [ADR-0007](docs/decisions/0007-wrap-openbot-keep-intelligence.md) says why, and
-onboarding says so before you type a key.*
+[](https://github.com/MasterYoav/xBot/actions/workflows/ci.yml)
+[](https://www.swift.org)
+[](https://www.apple.com/macos/)
+[](https://developer.apple.com/xcode/)
+[](docs/07-container-runtime.md)
+[](docs/04-model-providers.md)
+[](LICENSE)
@@ -19,12 +25,16 @@ onboarding says so before you type a key.*
## What it is
-xBot is a native macOS app. You download a `.dmg`, drag it to Applications, and open it. It walks
-you through everything else.
+xBot is a native macOS app. You download a `.dmg`, drag it to Applications, and open it. Five
+onboarding steps take you from there to your first agent: a system check, installing the container
+runtime for you if you have none, starting the engine, connecting a model, and meeting the agent.
-Behind the app, a full agent platform runs in containers on your Mac: each agent gets its own
+Behind the app, a full agent platform runs in a container on your Mac: each agent gets its own
computer with its own browser, its own files, and only the tools you grant it. Every action an agent
-takes is decided against a policy before it happens and recorded after.
+takes is decided against a policy before it happens and recorded in an append-only audit trail
+after. Model keys live in the macOS Keychain, and the engine's ports are bound to loopback only.
+Conversations are kept by the engine itself, in its own database, with no CopilotKit account
+([ADR-0008](docs/decisions/0008-local-thread-runner.md)).
You never open a terminal. You never edit a configuration file. You never read a log.
@@ -33,55 +43,68 @@ You never open a terminal. You never edit a configuration file. You never read a
Two good things existed separately.
[**OpenBot**](https://github.com/CopilotKit/openbot) is a serious, well-built, self-hosted agent
-platform — per-agent isolation, an action gateway, a real audit trail. It is also a developer
+platform, with per-agent isolation, an action gateway and a real audit trail. It is also a developer
template: you clone a repository, copy an `.env`, fill in credentials, and run a shell script.
-**Grok Bot** showed what the consumer shape of this looks like — a chat app with a rail of agents,
+**Grok Bot** showed what the consumer shape of this looks like: a chat app with a rail of agents,
a live view of what each one is doing, and settings you can actually find.
-xBot is the fusion: OpenBot's engine, a native Mac experience, and no lock-in to any single model
-vendor.
+xBot is the fusion: OpenBot's engine, wrapped rather than rewritten
+([ADR-0007](docs/decisions/0007-wrap-openbot-keep-intelligence.md)), a native Mac experience, and no
+lock-in to any single model vendor.
## Status
-**In development.** The native Mac client ships rail, conversation, composer, panel, command palette,
-onboarding (five steps), in-window settings (General, Models, Agents, Computer, Usage, Updates,
-Advanced), agent settings (model picker,
-plugins reach, handoff grants), plugins admin webview — all wired to `RuntimeController` and
-`HTTPEngineClient` when the engine is running. The app pulls a pinned ghcr engine digest from
-`manifests/engine-stable.json` on start. M2's model router has been driven live against a real vendor —
-per-run selection, a deployment fallback, an `openai-compatible` endpoint in the same process, and a
-bogus model name rejected by the vendor rather than silently substituted — and the hop the product
-actually uses, through `copilot.ts` and the AG-UI client, is covered by a test that asserts on the
-posted body. A second live vendor, M6 VM validation, and M7 signing/notarization remain open.
+**In development.** The native client ships the rail, conversation, composer, panel, command
+palette, onboarding, in-window settings (General, Models, Agents, Computer, Usage, Audit, Updates,
+Advanced), per-agent settings (model picker, plugin reach, handoff grants) and the plugins admin
+webview, all wired to the real engine. On start the app pulls the multi-arch engine image pinned by
+digest in [`manifests/engine-stable.json`](manifests/engine-stable.json).
+
+What has been proven against a real engine: the model router answered a real vendor (Anthropic)
+with per-run model selection, a deployment fallback and an `openai-compatible` endpoint in one
+process. On 27 September a throwaway engine with no CopilotKit key held a three-turn conversation
+with a local Ollama model, drove its browser through the client-tool loop, remembered across turns,
+and kept every conversation intact through a restart.
-Start at [`docs/README.md`](docs/README.md). The current milestone table is in
-[`docs/12-roadmap.md`](docs/12-roadmap.md).
+Still open: a second live vendor, one intermittent fault seen only under the full live suite, a
+clean-VM first run, and Developer ID signing, notarization, Sparkle keys and the first published
+release. Those last items need an account holder.
+
+Start at [`docs/README.md`](docs/README.md). The milestone table is in
+[`docs/12-roadmap.md`](docs/12-roadmap.md), and what is left, in order, is in
+[`docs/13-launch-checklist.md`](docs/13-launch-checklist.md).
### Run the Mac app locally
```sh
-cd apps/mac && swift run # debug: stub engine, full UI, no Docker
-cd apps/mac && XBOT_USE_RUNTIME=1 swift run # debug: real runtime path — Start in the UI
-cd apps/mac && swift test # 235 unit tests (SwiftPM)
-scripts/build-engine-image.sh # dev: build xbot/engine:1 for the runtime path
-scripts/check-engine-health.sh # dev: read-only /health check once the engine is up
-scripts/generate-app-icon.sh # compile xBot.icon → Assets.car + xBot.icns
-scripts/bundle-mac-app.sh # wrap release binary in XBot.app (after swift build -c release)
+scripts/generate-app-icon.sh # compile xBot.icon → Assets.car + xBot.icns (fresh clones; needs Xcode 26)
+cd apps/mac && swift run # debug: stub engine, full UI, no Docker
+cd apps/mac && XBOT_USE_RUNTIME=1 swift run # debug: real runtime path; press Start in the UI
+cd apps/mac && swift build --build-tests # compile the test targets too; plain `swift build` skips them
+cd apps/mac && swift test # 328 Swift tests; the live-engine suites skip without an engine
+scripts/build-engine-image.sh # dev: build xbot/engine:1 for the runtime path
+scripts/check-engine-health.sh # dev: read-only /health check once the engine is up
+eval "$(scripts/dev-db.sh)" # dev: pgvector database on port 55432 for the engine tests
+cd engine && bun run test:ci # engine tests, with a test-count floor
+scripts/bundle-mac-app.sh # wrap the release binary in XBot.app (after swift build -c release)
```
-Release builds always use the runtime path. On start the app fetches the pinned engine manifest
-(`manifests/engine-stable.json`) and pulls `ghcr.io/masteryoav/xbot-engine@sha256:…`. For local
-development, build `xbot/engine:1` with `scripts/build-engine-image.sh` or set
-`XBOT_ENGINE_IMAGE=xbot/engine:1`. Bearer token and encryption key are
-generated on first run and held in the Keychain. First Start can take up to ~2 minutes while
-Postgres initializes. If a start fails mid-boot, `docker rm -f xbot-engine` clears the container
-for a clean retry (volumes are kept).
+Requires macOS 14 or later and Xcode 26: the app icon is Icon Composer's format, and
+`Package.swift` expects the compiled icon, so `swift build` fails with "missing inputs" until
+`generate-app-icon.sh` has run. Release builds always use the runtime path and pull
+`ghcr.io/masteryoav/xbot-engine@sha256:…`; for local development build `xbot/engine:1` or set
+`XBOT_ENGINE_IMAGE=xbot/engine:1`. The bearer token and encryption key are generated on first run
+and held in the Keychain. The first Start can take up to about two minutes while Postgres
+initializes. If a start fails mid-boot, `docker rm -f xbot-engine` clears the container for a clean
+retry; the data volume is kept. Unsigned local builds are ad-hoc signed, so macOS may ask for
+Keychain access again after each rebuild.
## Documentation
| | |
| --- | --- |
+| [Documentation index](docs/README.md) | Reading order and conventions |
| [Vision](docs/01-vision.md) | What we are building, for whom, and what we are not building |
| [Architecture](docs/02-architecture.md) | Services, ports, data flow |
| [The OpenBot fork](docs/03-openbot-fork.md) | What we inherit and what we have to change |
@@ -96,14 +119,22 @@ for a clean retry (volumes are kept).
| [Roadmap](docs/12-roadmap.md) | Milestones |
| [Launch checklist](docs/13-launch-checklist.md) | What is left between here and a download, in order |
| [Engine environment mapping](docs/env-mapping.md) | App settings → container env vars |
-| [Decisions](docs/decisions/) | ADRs — read these before disagreeing with anything above |
+| [Decisions](docs/decisions/) | ADRs 0001–0008. Read these before disagreeing with anything above |
+| [Plan: computer client tools](docs/plans/computer-client-tools.md) | Giving agents their browser, files and shell through the Mac client |
+| [Plan: managed Bot and model keys](docs/plans/managed-bot-and-model-keys.md) | A Bot in the engine image, and model keys that reach it |
+| [Phase 1: ship readiness](docs/superpowers/specs/2026-09-17-phase-1-ship-readiness-design.md) | The evidence needed before a stranger can install xBot |
## Built on OpenBot
xBot's engine is a fork of [OpenBot](https://github.com/CopilotKit/openbot) by
[CopilotKit](https://copilotkit.ai), used under the MIT licence. Copyright © 2026 CopilotKit.
-See [`NOTICE`](NOTICE).
+See [`NOTICE`](NOTICE) and [`engine/LICENSE`](engine/LICENSE). App updates use
+[Sparkle](https://sparkle-project.org), also MIT-licensed.
+
+xBot is not affiliated with, endorsed by, or sponsored by CopilotKit, xAI, X Corp., or any model
+provider.
## Licence
-MIT.
+MIT. See [`LICENSE`](LICENSE). The engine keeps OpenBot's own MIT licence in
+[`engine/LICENSE`](engine/LICENSE).
diff --git a/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift b/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift
index caed25a..1adcb8d 100644
--- a/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift
+++ b/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift
@@ -70,23 +70,25 @@ struct LiveEngineTests {
the vendor the failure would be the router's "no key" sentence, or no managed Bot at all. Only a
key that travelled the whole way comes back as Anthropic refusing it.
*/
+ // OpenAI, not the Anthropic key the next test stores: Swift Testing runs them at once, and two
+ // replaces of one key racing each other is a 500 from the engine's vault, not a client bug.
@Test func aStoredKeyReachesTheVault() async throws {
let client = client
try await client.storeModelKey(
- "sk-ant-api03-xbot-live-check-deliberately-invalid",
- providerId: "anthropic", baseURL: nil, fingerprint: "live-check"
+ "sk-xbot-live-check-deliberately-invalid",
+ providerId: "openai", baseURL: nil, fingerprint: "live-check"
)
- let stored = try await client.liveModelKeys().first { $0.keyId == "xbot-model:anthropic" }
+ let stored = try await client.liveModelKeys().first { $0.keyId == "xbot-model:openai" }
#expect(stored?.fingerprint == "live-check")
// Replacing it leaves one live key, not two.
- try await client.storeModelKey("sk-ant-replaced", providerId: "anthropic", baseURL: nil, fingerprint: "live-check-2")
- let live = try await client.liveModelKeys().filter { $0.keyId == "xbot-model:anthropic" }
+ try await client.storeModelKey("sk-replaced", providerId: "openai", baseURL: nil, fingerprint: "live-check-2")
+ let live = try await client.liveModelKeys().filter { $0.keyId == "xbot-model:openai" }
#expect(live.count == 1)
#expect(live.first?.fingerprint == "live-check-2")
if let id = live.first?.id { try await client.revokeModelKey(credentialId: id) }
- #expect(try await client.liveModelKeys().allSatisfy { $0.keyId != "xbot-model:anthropic" })
+ #expect(try await client.liveModelKeys().allSatisfy { $0.keyId != "xbot-model:openai" })
}
/// Since ADR-0008, `LocalThreadRunner` answers this without an Intelligence key — the throwaway
@@ -248,3 +250,90 @@ struct LiveComputerToolsTests {
#expect(refused != nil)
}
}
+
+/**
+ A conversation that succeeds, against a real model — a local Ollama, so no key is needed.
+
+ Skipped unless `XBOT_LIVE_OLLAMA_MODEL` is set as well, to a pulled model that can call tools
+ (`qwen2.5:7b` works). Everything above proves how a turn fails; this is the one place a turn is
+ proven to answer, remember within a conversation, drive its computer through the client-tool loop,
+ and leave the thread holding each message once. Launch checklist item 5, the engine half.
+
+ A model is not deterministic, so the prompts ask for things a small one reliably gets right and the
+ assertions check for a word, never a sentence.
+ */
+@Suite(.enabled(if: ProcessInfo.processInfo.environment["XBOT_LIVE_ENGINE_URL"] != nil
+ && ProcessInfo.processInfo.environment["XBOT_LIVE_OLLAMA_MODEL"] != nil))
+struct LiveConversationTests {
+ private let env = ProcessInfo.processInfo.environment
+
+ private var client: HTTPEngineClient {
+ let url = URL(string: env["XBOT_LIVE_ENGINE_URL"] ?? "http://127.0.0.1:3001")!
+ let token = env["XBOT_LIVE_ENGINE_TOKEN_FILE"].flatMap {
+ try? String(contentsOfFile: $0, encoding: .utf8).trimmingCharacters(in: .whitespacesAndNewlines)
+ }
+ return HTTPEngineClient(baseURL: url, token: token)
+ }
+
+ private struct Turn {
+ var text = ""
+ var tools: [String] = []
+ var failure: String?
+ }
+
+ private func turn(_ text: String, in channel: Channel.ID) async throws -> Turn {
+ var turn = Turn()
+ for try await event in client.send(text, to: channel) {
+ switch event {
+ case .textDelta(_, let delta): turn.text += delta
+ case .toolCall(_, _, let name, _): turn.tools.append(name)
+ case .failed(_, let reason): turn.failure = reason
+ default: continue
+ }
+ }
+ print("live turn:", text, "→", turn.tools, turn.failure ?? turn.text)
+ return turn
+ }
+
+ @Test func aConversationAnswersRemembersUsesItsComputerAndIsKeptOnce() async throws {
+ let client = client
+ // What `ModelProviderCatalog.engineRouting(for: "ollama")` sends, which lives in XBotCore.
+ let agent = try await client.createAgent(AgentDraft(
+ name: "Conversation check",
+ model: ModelSelection(
+ provider: "Ollama", providerID: "openai-compatible", model: env["XBOT_LIVE_OLLAMA_MODEL"]!,
+ baseURL: env["XBOT_LIVE_OLLAMA_BASE_URL"] ?? "http://host.docker.internal:11434/v1", capabilities: []
+ )
+ ))
+ let channel = try await client.createChannel(agentIds: [agent.id])
+
+ let told = try await turn("Remember this codeword: tangerine. Reply with just OK.", in: channel.id)
+ #expect(told.failure == nil)
+ #expect(!told.text.isEmpty)
+
+ // A 7B model sometimes writes the call out as text instead of making it. That is the model's
+ // choice, not the loop failing, so it gets one more ask before the loop is judged.
+ var sent = 2
+ var browsed = try await turn(
+ "Use your browser to open https://example.com, then tell me the page's title.", in: channel.id
+ )
+ if !browsed.tools.contains("computer_navigate") {
+ sent += 1
+ browsed = try await turn("Call the computer_navigate tool with https://example.com now.", in: channel.id)
+ }
+ #expect(browsed.failure == nil)
+ #expect(browsed.tools.contains("computer_navigate"))
+ #expect(browsed.text.localizedCaseInsensitiveContains("example"))
+
+ // Two turns and a tool loop later. Each run carries the whole conversation, so this is memory.
+ let asked = try await turn("What was the codeword I gave you? Answer with the word only.", in: channel.id)
+ #expect(asked.failure == nil)
+ #expect(asked.text.localizedCaseInsensitiveContains("tangerine"))
+
+ // The tool loop resends messages the thread already holds and relies on their ids matching.
+ let history = try await client.messages(in: channel.id)
+ print("live history:", history.map { "\($0.isFromUser ? "user" : "agent") \($0.id) \($0.text.prefix(40))" })
+ #expect(Set(history.map(\.id)).count == history.count)
+ #expect(history.filter(\.isFromUser).count == sent + 1)
+ }
+}
diff --git a/docs/12-roadmap.md b/docs/12-roadmap.md
index cae308c..0ca6108 100644
--- a/docs/12-roadmap.md
+++ b/docs/12-roadmap.md
@@ -141,7 +141,9 @@ and `openai-compatible` — rather than two vendors.
An end-to-end conversation through the server no longer needs Intelligence credentials: with none
of the `INTELLIGENCE_*` variables set, `copilot.ts` runs on `LocalThreadRunner` (ADR-0008). That run
-through the server is launch checklist item 5, and has not been done yet.
+through the server is launch checklist item 5. Its engine half was run on 27 September against a
+local Ollama model, and found three faults on this hop that every conversation hit; all three are
+fixed and recorded there.
**Still open:** a second live vendor. The three native adapters are built and each is covered by a
test asserting which client a selection gets (`agent-langgraph/tests/models-build.test.ts`). Usage accounting is done — the agent sums `usage_metadata` across a turn and emits
diff --git a/docs/13-launch-checklist.md b/docs/13-launch-checklist.md
index 66e3885..c69a60c 100644
--- a/docs/13-launch-checklist.md
+++ b/docs/13-launch-checklist.md
@@ -134,11 +134,33 @@ engine (`docs/plans/computer-client-tools.md`):
**There is an automated version of the engine half.** `apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift`
drives the real client against a running engine — point it at a throwaway one, since it creates
-agents. Its conversation-store assertions now run against local mode by default; a separate
-`XBOT_LIVE_ENGINE_HAS_INTELLIGENCE=1` branch exists only for a deployment that configures
-Intelligence by hand, which nothing here needs. This suite, and `LocalThreadRunner`'s own
-integration test (`engine/server/tests/local-thread-runner.integration.test.ts`), still need to be
-run before this item can be checked off — see the verification commands in `CLAUDE.md`.
+agents. **It has now been run** (27 September), against an engine built from the working tree and a
+local Ollama (`qwen2.5:7b`): set `XBOT_LIVE_OLLAMA_MODEL` and `LiveConversationTests` holds a real
+three-turn conversation — it answers, drives `computer_navigate` through the client-tool loop and
+reports the page, remembers a codeword two turns later, and leaves each message in the thread once.
+Restarting that engine left every conversation's message count unchanged. What remains for a person
+is the app half: the same conversation typed into the window, and the asks.
+
+That first run found three engine faults, each of which failed every conversation:
+
+- **The managed Bot could not be dialled.** `createAgentFetch` refused `127.0.0.1:4201`, the
+ address the image itself gives the Bot, so every run through the server ended in "This deployment
+ will not dial…". The Bot's configured host is now added to the allowed hosts
+ (`server/src/agents/managed-agent-host.ts`), and only that host and port.
+- **A keyless `openai-compatible` endpoint was refused locally.** OpenAI's client falls back to
+ `OPENAI_API_KEY` when handed no key, so every Ollama run failed with "Missing credentials". It now
+ gets a placeholder the endpoint ignores (`agent-langgraph/src/models/build.ts`).
+- **Every Anthropic run was refused.** The server's standing role message arrived after the Bot's
+ guidance, and Anthropic accepts a system message only first. All system texts are now merged, in
+ order, into one leading message (`agent-langgraph/src/history.ts`).
+
+**One open finding.** On a freshly started engine, roughly one full live run in six had a turn come
+back empty, or a request go unanswered, while all nine live tests ran at once. The engine logged
+"Thread already running" or, once, "No agents are registered", both of which should be impossible
+there. It was never seen with the conversation suite alone (5 of 5 fresh engines), nor with a
+logging proxy in the path (12 of 12). A leftover process, a database lock, identity, idle
+keep-alive and split request writes were each checked and ruled out. Worth one more look before
+launch, starting from the app under ordinary use rather than nine tests at once.
**Also close the last M2 item here:** point one agent at a second real vendor — Anthropic is already
proven, so use OpenAI or Google — and confirm the reply comes from the vendor you picked. A model
@@ -197,6 +219,10 @@ From a shell holding no credentials:
passed all eight live client tests: health, read endpoints, the vault, agent and conversation
round-trip, browser, files, shell, and the asks. That run predates ADR-0008; the conversation store
now needs re-running against `LocalThreadRunner` rather than a CopilotKit key — see item 5.
+- On 27 September, it was: a throwaway engine in local mode with no CopilotKit key held a real
+ conversation with a local Ollama model, used its computer through the client-tool loop, remembered
+ across turns, and kept every conversation intact through a restart. Item 5 has the details and
+ the one open finding.
- The pinned engine is one multi-arch image (linux/amd64 and linux/arm64). On 14 September an Apple
Silicon Mac pulled it anonymously, got arm64, was healthy in about ten seconds on a port other than
3001, and passed the live suite: browser, files, shell, and a help request handed back. Before that
diff --git a/docs/plans/computer-client-tools.md b/docs/plans/computer-client-tools.md
index 07a6a93..1c292dd 100644
--- a/docs/plans/computer-client-tools.md
+++ b/docs/plans/computer-client-tools.md
@@ -32,11 +32,14 @@ been offered a browser, files, or a shell — the product's headline feature doe
- **C1 done**, verified against a real engine: navigate, snapshot, write/read a file, run a command
(`LiveComputerToolsTests`).
- **C2 done.** Every run carries `ComputerTools.wireTools`.
-- **C3 done, verified only against a stubbed engine.** `WireTranscript` rebuilds the conversation with
+- **C3 done, verified against a real engine** on 27 September (`LiveConversationTests`, local
+ Ollama, no CopilotKit key): the follow-up run continues after `computer_navigate`, and the thread
+ holds each message once.
+ `WireTranscript` rebuilds the conversation with
`@ag-ui/client`'s reducer, id for id, starting from the thread's stored history — which also fixes
agents having no memory of earlier messages. `HTTPEngineClient.stream` runs pending calls at
- `RUN_FINISHED` and follows up, capped at 25 rounds. **Needs one real conversation with a CopilotKit
- key** before v1: that the history route's rows round-trip, and that follow-ups do not duplicate.
+ `RUN_FINISHED` and follows up, capped at 25 rounds. The real conversation this waited on no longer
+ needs a CopilotKit key (ADR-0008), and has been held: rows round-trip and follow-ups do not duplicate.
- **C4 done, verified against a stubbed engine.** While a turn runs, `AppState` polls `/control`
every two seconds, as upstream's client does. A help request shows "needs you" with Take control,
which opens the screen; handing back ends the wait. A secret request shows a masked field that posts
diff --git a/engine/agent-langgraph/src/history.ts b/engine/agent-langgraph/src/history.ts
index 2f98252..75a5132 100644
--- a/engine/agent-langgraph/src/history.ts
+++ b/engine/agent-langgraph/src/history.ts
@@ -24,6 +24,14 @@ export { NO_ANSWER_CAME };
/** Translate the conversation AG-UI carries into LangChain's message classes. */
export function toLangChainMessages(input: RunAgentInput): BaseMessage[] {
const messages: BaseMessage[] = [new SystemMessage(COMPUTER_GUIDANCE)];
+ /*
+ * Every system text, gathered into that first message rather than pushed where it arrived.
+ *
+ * The server sends standing messages of its own — the role, what this Bot holds. OpenAI answers
+ * system messages anywhere; Anthropic refuses the whole run: "System messages are only permitted
+ * as the first passed message". Merged in order, so what each one says is unchanged.
+ */
+ const system = [COMPUTER_GUIDANCE];
/*
* Which calls in this history were ever answered.
@@ -53,7 +61,7 @@ export function toLangChainMessages(input: RunAgentInput): BaseMessage[] {
continue;
}
if (message.role === "system" || message.role === "developer") {
- messages.push(new SystemMessage(String(message.content ?? "")));
+ system.push(String(message.content ?? ""));
continue;
}
if (message.role === "tool") {
@@ -123,6 +131,7 @@ export function toLangChainMessages(input: RunAgentInput): BaseMessage[] {
messages.push(new HumanMessage(CONTINUE_TURN));
}
+ messages[0] = new SystemMessage(system.join("\n\n"));
return messages;
}
diff --git a/engine/agent-langgraph/src/models/build.ts b/engine/agent-langgraph/src/models/build.ts
index 23be3fd..510e59a 100644
--- a/engine/agent-langgraph/src/models/build.ts
+++ b/engine/agent-langgraph/src/models/build.ts
@@ -80,7 +80,11 @@ export function buildChatModel({
*/
return new ChatOpenAI({
model: resolved.model,
- apiKey: resolved.apiKey,
+ // A keyless compatible endpoint (a local Ollama) still needs *a* key: without one the client
+ // falls back to OPENAI_API_KEY and refuses the run. The endpoint ignores whatever is sent.
+ apiKey:
+ resolved.apiKey ||
+ (resolved.providerId === "openai-compatible" ? "not-needed" : undefined),
streaming: true,
...(resolved.baseURL
? { configuration: { baseURL: resolved.baseURL } }
diff --git a/engine/agent-langgraph/tests/history-system.test.ts b/engine/agent-langgraph/tests/history-system.test.ts
new file mode 100644
index 0000000..b56bc6c
--- /dev/null
+++ b/engine/agent-langgraph/tests/history-system.test.ts
@@ -0,0 +1,38 @@
+import { describe, expect, test } from "bun:test";
+import type { RunAgentInput } from "@ag-ui/core";
+import { SystemMessage } from "@langchain/core/messages";
+import { COMPUTER_GUIDANCE } from "../../shared/bot-prompt";
+import { toLangChainMessages } from "../src/history";
+
+/**
+ * Every system text arrives as one leading system message.
+ *
+ * The server sends its own standing messages — the role, what this Bot holds — and the Bot puts its
+ * guidance first. OpenAI answers any number of system messages anywhere; Anthropic refuses the run
+ * outright ("System messages are only permitted as the first passed message"), so every Anthropic
+ * agent failed before reaching the vendor. Found by the first live run through the server.
+ */
+const input = (messages: unknown[]): RunAgentInput =>
+ ({ messages }) as unknown as RunAgentInput;
+
+describe("system messages", () => {
+ test("are merged, in order, into the first message and nowhere else", () => {
+ const messages = toLangChainMessages(
+ input([
+ { role: "system", content: "You are the research agent." },
+ { role: "user", content: "hello" },
+ { role: "developer", content: "You hold Drive." },
+ ]),
+ );
+
+ const system = messages.filter((m) => m instanceof SystemMessage);
+ expect(system).toHaveLength(1);
+ expect(messages[0]).toBe(system[0]);
+ const text = String(system[0]?.content);
+ expect(text.indexOf(COMPUTER_GUIDANCE)).toBe(0);
+ expect(text.indexOf("You are the research agent.")).toBeLessThan(
+ text.indexOf("You hold Drive."),
+ );
+ expect(messages).toHaveLength(2);
+ });
+});
diff --git a/engine/agent-langgraph/tests/models-build.test.ts b/engine/agent-langgraph/tests/models-build.test.ts
index d5b1757..ef427da 100644
--- a/engine/agent-langgraph/tests/models-build.test.ts
+++ b/engine/agent-langgraph/tests/models-build.test.ts
@@ -54,6 +54,21 @@ describe("the client a selection is answered by", () => {
);
});
+ test("a keyless openai-compatible endpoint still hands the client a key", () => {
+ // A local Ollama asks for none, but OpenAI's client falls back to OPENAI_API_KEY and refuses the
+ // run without one: "Missing credentials", on every turn, found by the first live Ollama run.
+ const model = buildChatModel({
+ selection: {
+ providerId: "openai-compatible",
+ model: "qwen2.5:7b",
+ baseURL: "http://host.docker.internal:11434/v1",
+ },
+ fallback: undefined,
+ keys: {},
+ }) as ChatOpenAI;
+ expect(model.apiKey).toBeTruthy();
+ });
+
test("an agent that never chose inherits the workspace default", () => {
const model = buildChatModel({
selection: undefined,
diff --git a/engine/server/src/agents/managed-agent-host.ts b/engine/server/src/agents/managed-agent-host.ts
new file mode 100644
index 0000000..d19f119
--- /dev/null
+++ b/engine/server/src/agents/managed-agent-host.ts
@@ -0,0 +1,20 @@
+/**
+ * The addresses every run may dial: the ones the deployment named, plus its own Bot's.
+ *
+ * The one-container image runs the Bot beside the API at a loopback address, which the
+ * private-address guard in `endpoint.ts` refuses — correctly, for an address a person typed. This
+ * one came from the deployment's own configuration (`MANAGED_AGENT_AG_UI_URL`), which is exactly
+ * what `allowedHosts` exists for. Host and port together, so the rest of loopback — the database,
+ * the computer — stays refused. Compose never needed this: there the Bot is a service name, which is
+ * not a private literal.
+ *
+ * Dialling only. Registration keeps the named list alone, so nobody can register an agent at the
+ * managed Bot's address by way of this.
+ */
+export function dialableHosts(
+ named: ReadonlySet,
+ managedAgent: { endpoint: URL } | undefined,
+): ReadonlySet {
+ if (!managedAgent) return named;
+ return new Set([...named, managedAgent.endpoint.host]);
+}
diff --git a/engine/server/src/index.ts b/engine/server/src/index.ts
index d7f31d4..eb6af4c 100644
--- a/engine/server/src/index.ts
+++ b/engine/server/src/index.ts
@@ -8,6 +8,7 @@ import { eq } from "drizzle-orm";
import { COMPUTER_GUIDANCE } from "../../shared/bot-prompt";
import { mintRunAssertion, readRunAssertion } from "./agents/callback-token";
import { createAgentFetch } from "./agents/endpoint";
+import { dialableHosts } from "./agents/managed-agent-host";
import { askTheirOwnPerson, escalationTool } from "./agents/escalation";
import { createHandoffDesk, HANDOFF_KIND } from "./agents/handoff";
import { createHandoffDelivery } from "./agents/handoff-delivery";
@@ -603,8 +604,12 @@ const selectionForActor = (actorId: string): ToolSelection => ({
// reading and the same one `createApp` takes.
const agentFetch = createAgentFetch({
allowPrivateHosts: config.computer?.allowPrivateHosts === true,
- // Named addresses are reachable on every hop, not only the one that was registered.
- allowedHosts: config.agentEndpointAllowedHosts,
+ // Named addresses are reachable on every hop, not only the one that was registered. The
+ // deployment's own Bot is one of them: in the one-container image it sits on loopback.
+ allowedHosts: dialableHosts(
+ config.agentEndpointAllowedHosts,
+ config.managedAgent,
+ ),
// The refusal is what the run already knows; this is what the deployment knows. Written here
// rather than in `endpoint.ts` so that file keeps deciding and nothing else, the way the
// target check it reuses does.
diff --git a/engine/server/tests/managed-agent-dial.test.ts b/engine/server/tests/managed-agent-dial.test.ts
new file mode 100644
index 0000000..b9613ea
--- /dev/null
+++ b/engine/server/tests/managed-agent-dial.test.ts
@@ -0,0 +1,40 @@
+import { describe, expect, test } from "bun:test";
+import { createAgentFetch } from "../src/agents/endpoint";
+import { dialableHosts } from "../src/agents/managed-agent-host";
+
+// The one-container image runs the Bot beside the API at a loopback address (docker/s6 api/run).
+// Every run dials it through the private-address guard, which refused it: no conversation through
+// the server could ever be answered.
+const managed = {
+ endpoint: new URL("http://127.0.0.1:4201/ag-ui"),
+ token: "t",
+};
+const answering = (async () =>
+ new Response("ok", { status: 200 })) as unknown as typeof fetch;
+
+describe("the deployment's own Bot", () => {
+ test("is dialled", async () => {
+ const dial = createAgentFetch({
+ allowedHosts: dialableHosts(new Set(), managed),
+ fetchImpl: answering,
+ });
+ expect((await dial("http://127.0.0.1:4201/ag-ui")).status).toBe(200);
+ });
+
+ test("names its port, so the rest of loopback stays refused", async () => {
+ const dial = createAgentFetch({
+ allowedHosts: dialableHosts(new Set(), managed),
+ fetchImpl: answering,
+ });
+ await expect(dial("http://127.0.0.1:5432/")).rejects.toThrow();
+ });
+
+ test("adds nothing when there is none, and keeps what was named", () => {
+ const named = new Set(["agents.internal"]);
+ expect(dialableHosts(named, undefined)).toBe(named);
+ expect([...dialableHosts(named, managed)].sort()).toEqual([
+ "127.0.0.1:4201",
+ "agents.internal",
+ ]);
+ });
+});