diff --git a/LICENSE b/LICENSE new file mode 100644 index 0000000..7c254dc --- /dev/null +++ b/LICENSE @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 xBot contributors + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/README.md b/README.md index dbbdaf2..d154fcd 100644 --- a/README.md +++ b/README.md @@ -2,16 +2,22 @@ xBot - -**Your own AI coworkers, on your own Mac.** +

Your own AI coworkers, on your own Mac

Create agents, give them a computer, watch them work, and take the wheel when you want to. -Bring any model — OpenAI, Anthropic, Google, xAI, or a model running locally through Ollama. -Your agents, their files, and their browsers stay on this Mac. +Bring any model: OpenAI, Anthropic, Google, xAI, or a model running locally through Ollama. +Your agents, their files, their browsers and your conversations stay on this Mac. + +*Not released yet: there is no signed download. Build it from source below, and see the +[launch checklist](docs/13-launch-checklist.md) for what stands between here and a `.dmg`.* -*In v1 your conversation history is the exception: it is stored by CopilotKit, the service xBot's -engine is built on. [ADR-0007](docs/decisions/0007-wrap-openbot-keep-intelligence.md) says why, and -onboarding says so before you type a key.* +[![CI](https://github.com/MasterYoav/xBot/actions/workflows/ci.yml/badge.svg)](https://github.com/MasterYoav/xBot/actions/workflows/ci.yml) +[![Swift](https://img.shields.io/badge/Swift-F54A2A?logo=swift&logoColor=white)](https://www.swift.org) +[![macOS](https://img.shields.io/badge/macOS-000000?logo=apple&logoColor=F0F0F0)](https://www.apple.com/macos/) +[![Xcode](https://img.shields.io/badge/Xcode-007ACC?logo=Xcode&logoColor=white)](https://developer.apple.com/xcode/) +[![Docker](https://img.shields.io/badge/Docker-2496ED?logo=docker&logoColor=fff)](docs/07-container-runtime.md) +[![Ollama](https://img.shields.io/badge/Ollama-fff?logo=ollama&logoColor=000)](docs/04-model-providers.md) +[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE) @@ -19,12 +25,16 @@ onboarding says so before you type a key.* ## What it is -xBot is a native macOS app. You download a `.dmg`, drag it to Applications, and open it. It walks -you through everything else. +xBot is a native macOS app. You download a `.dmg`, drag it to Applications, and open it. Five +onboarding steps take you from there to your first agent: a system check, installing the container +runtime for you if you have none, starting the engine, connecting a model, and meeting the agent. -Behind the app, a full agent platform runs in containers on your Mac: each agent gets its own +Behind the app, a full agent platform runs in a container on your Mac: each agent gets its own computer with its own browser, its own files, and only the tools you grant it. Every action an agent -takes is decided against a policy before it happens and recorded after. +takes is decided against a policy before it happens and recorded in an append-only audit trail +after. Model keys live in the macOS Keychain, and the engine's ports are bound to loopback only. +Conversations are kept by the engine itself, in its own database, with no CopilotKit account +([ADR-0008](docs/decisions/0008-local-thread-runner.md)). You never open a terminal. You never edit a configuration file. You never read a log. @@ -33,55 +43,68 @@ You never open a terminal. You never edit a configuration file. You never read a Two good things existed separately. [**OpenBot**](https://github.com/CopilotKit/openbot) is a serious, well-built, self-hosted agent -platform — per-agent isolation, an action gateway, a real audit trail. It is also a developer +platform, with per-agent isolation, an action gateway and a real audit trail. It is also a developer template: you clone a repository, copy an `.env`, fill in credentials, and run a shell script. -**Grok Bot** showed what the consumer shape of this looks like — a chat app with a rail of agents, +**Grok Bot** showed what the consumer shape of this looks like: a chat app with a rail of agents, a live view of what each one is doing, and settings you can actually find. -xBot is the fusion: OpenBot's engine, a native Mac experience, and no lock-in to any single model -vendor. +xBot is the fusion: OpenBot's engine, wrapped rather than rewritten +([ADR-0007](docs/decisions/0007-wrap-openbot-keep-intelligence.md)), a native Mac experience, and no +lock-in to any single model vendor. ## Status -**In development.** The native Mac client ships rail, conversation, composer, panel, command palette, -onboarding (five steps), in-window settings (General, Models, Agents, Computer, Usage, Updates, -Advanced), agent settings (model picker, -plugins reach, handoff grants), plugins admin webview — all wired to `RuntimeController` and -`HTTPEngineClient` when the engine is running. The app pulls a pinned ghcr engine digest from -`manifests/engine-stable.json` on start. M2's model router has been driven live against a real vendor — -per-run selection, a deployment fallback, an `openai-compatible` endpoint in the same process, and a -bogus model name rejected by the vendor rather than silently substituted — and the hop the product -actually uses, through `copilot.ts` and the AG-UI client, is covered by a test that asserts on the -posted body. A second live vendor, M6 VM validation, and M7 signing/notarization remain open. +**In development.** The native client ships the rail, conversation, composer, panel, command +palette, onboarding, in-window settings (General, Models, Agents, Computer, Usage, Audit, Updates, +Advanced), per-agent settings (model picker, plugin reach, handoff grants) and the plugins admin +webview, all wired to the real engine. On start the app pulls the multi-arch engine image pinned by +digest in [`manifests/engine-stable.json`](manifests/engine-stable.json). + +What has been proven against a real engine: the model router answered a real vendor (Anthropic) +with per-run model selection, a deployment fallback and an `openai-compatible` endpoint in one +process. On 27 September a throwaway engine with no CopilotKit key held a three-turn conversation +with a local Ollama model, drove its browser through the client-tool loop, remembered across turns, +and kept every conversation intact through a restart. -Start at [`docs/README.md`](docs/README.md). The current milestone table is in -[`docs/12-roadmap.md`](docs/12-roadmap.md). +Still open: a second live vendor, one intermittent fault seen only under the full live suite, a +clean-VM first run, and Developer ID signing, notarization, Sparkle keys and the first published +release. Those last items need an account holder. + +Start at [`docs/README.md`](docs/README.md). The milestone table is in +[`docs/12-roadmap.md`](docs/12-roadmap.md), and what is left, in order, is in +[`docs/13-launch-checklist.md`](docs/13-launch-checklist.md). ### Run the Mac app locally ```sh -cd apps/mac && swift run # debug: stub engine, full UI, no Docker -cd apps/mac && XBOT_USE_RUNTIME=1 swift run # debug: real runtime path — Start in the UI -cd apps/mac && swift test # 235 unit tests (SwiftPM) -scripts/build-engine-image.sh # dev: build xbot/engine:1 for the runtime path -scripts/check-engine-health.sh # dev: read-only /health check once the engine is up -scripts/generate-app-icon.sh # compile xBot.icon → Assets.car + xBot.icns -scripts/bundle-mac-app.sh # wrap release binary in XBot.app (after swift build -c release) +scripts/generate-app-icon.sh # compile xBot.icon → Assets.car + xBot.icns (fresh clones; needs Xcode 26) +cd apps/mac && swift run # debug: stub engine, full UI, no Docker +cd apps/mac && XBOT_USE_RUNTIME=1 swift run # debug: real runtime path; press Start in the UI +cd apps/mac && swift build --build-tests # compile the test targets too; plain `swift build` skips them +cd apps/mac && swift test # 328 Swift tests; the live-engine suites skip without an engine +scripts/build-engine-image.sh # dev: build xbot/engine:1 for the runtime path +scripts/check-engine-health.sh # dev: read-only /health check once the engine is up +eval "$(scripts/dev-db.sh)" # dev: pgvector database on port 55432 for the engine tests +cd engine && bun run test:ci # engine tests, with a test-count floor +scripts/bundle-mac-app.sh # wrap the release binary in XBot.app (after swift build -c release) ``` -Release builds always use the runtime path. On start the app fetches the pinned engine manifest -(`manifests/engine-stable.json`) and pulls `ghcr.io/masteryoav/xbot-engine@sha256:…`. For local -development, build `xbot/engine:1` with `scripts/build-engine-image.sh` or set -`XBOT_ENGINE_IMAGE=xbot/engine:1`. Bearer token and encryption key are -generated on first run and held in the Keychain. First Start can take up to ~2 minutes while -Postgres initializes. If a start fails mid-boot, `docker rm -f xbot-engine` clears the container -for a clean retry (volumes are kept). +Requires macOS 14 or later and Xcode 26: the app icon is Icon Composer's format, and +`Package.swift` expects the compiled icon, so `swift build` fails with "missing inputs" until +`generate-app-icon.sh` has run. Release builds always use the runtime path and pull +`ghcr.io/masteryoav/xbot-engine@sha256:…`; for local development build `xbot/engine:1` or set +`XBOT_ENGINE_IMAGE=xbot/engine:1`. The bearer token and encryption key are generated on first run +and held in the Keychain. The first Start can take up to about two minutes while Postgres +initializes. If a start fails mid-boot, `docker rm -f xbot-engine` clears the container for a clean +retry; the data volume is kept. Unsigned local builds are ad-hoc signed, so macOS may ask for +Keychain access again after each rebuild. ## Documentation | | | | --- | --- | +| [Documentation index](docs/README.md) | Reading order and conventions | | [Vision](docs/01-vision.md) | What we are building, for whom, and what we are not building | | [Architecture](docs/02-architecture.md) | Services, ports, data flow | | [The OpenBot fork](docs/03-openbot-fork.md) | What we inherit and what we have to change | @@ -96,14 +119,22 @@ for a clean retry (volumes are kept). | [Roadmap](docs/12-roadmap.md) | Milestones | | [Launch checklist](docs/13-launch-checklist.md) | What is left between here and a download, in order | | [Engine environment mapping](docs/env-mapping.md) | App settings → container env vars | -| [Decisions](docs/decisions/) | ADRs — read these before disagreeing with anything above | +| [Decisions](docs/decisions/) | ADRs 0001–0008. Read these before disagreeing with anything above | +| [Plan: computer client tools](docs/plans/computer-client-tools.md) | Giving agents their browser, files and shell through the Mac client | +| [Plan: managed Bot and model keys](docs/plans/managed-bot-and-model-keys.md) | A Bot in the engine image, and model keys that reach it | +| [Phase 1: ship readiness](docs/superpowers/specs/2026-09-17-phase-1-ship-readiness-design.md) | The evidence needed before a stranger can install xBot | ## Built on OpenBot xBot's engine is a fork of [OpenBot](https://github.com/CopilotKit/openbot) by [CopilotKit](https://copilotkit.ai), used under the MIT licence. Copyright © 2026 CopilotKit. -See [`NOTICE`](NOTICE). +See [`NOTICE`](NOTICE) and [`engine/LICENSE`](engine/LICENSE). App updates use +[Sparkle](https://sparkle-project.org), also MIT-licensed. + +xBot is not affiliated with, endorsed by, or sponsored by CopilotKit, xAI, X Corp., or any model +provider. ## Licence -MIT. +MIT. See [`LICENSE`](LICENSE). The engine keeps OpenBot's own MIT licence in +[`engine/LICENSE`](engine/LICENSE). diff --git a/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift b/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift index caed25a..1adcb8d 100644 --- a/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift +++ b/apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift @@ -70,23 +70,25 @@ struct LiveEngineTests { the vendor the failure would be the router's "no key" sentence, or no managed Bot at all. Only a key that travelled the whole way comes back as Anthropic refusing it. */ + // OpenAI, not the Anthropic key the next test stores: Swift Testing runs them at once, and two + // replaces of one key racing each other is a 500 from the engine's vault, not a client bug. @Test func aStoredKeyReachesTheVault() async throws { let client = client try await client.storeModelKey( - "sk-ant-api03-xbot-live-check-deliberately-invalid", - providerId: "anthropic", baseURL: nil, fingerprint: "live-check" + "sk-xbot-live-check-deliberately-invalid", + providerId: "openai", baseURL: nil, fingerprint: "live-check" ) - let stored = try await client.liveModelKeys().first { $0.keyId == "xbot-model:anthropic" } + let stored = try await client.liveModelKeys().first { $0.keyId == "xbot-model:openai" } #expect(stored?.fingerprint == "live-check") // Replacing it leaves one live key, not two. - try await client.storeModelKey("sk-ant-replaced", providerId: "anthropic", baseURL: nil, fingerprint: "live-check-2") - let live = try await client.liveModelKeys().filter { $0.keyId == "xbot-model:anthropic" } + try await client.storeModelKey("sk-replaced", providerId: "openai", baseURL: nil, fingerprint: "live-check-2") + let live = try await client.liveModelKeys().filter { $0.keyId == "xbot-model:openai" } #expect(live.count == 1) #expect(live.first?.fingerprint == "live-check-2") if let id = live.first?.id { try await client.revokeModelKey(credentialId: id) } - #expect(try await client.liveModelKeys().allSatisfy { $0.keyId != "xbot-model:anthropic" }) + #expect(try await client.liveModelKeys().allSatisfy { $0.keyId != "xbot-model:openai" }) } /// Since ADR-0008, `LocalThreadRunner` answers this without an Intelligence key — the throwaway @@ -248,3 +250,90 @@ struct LiveComputerToolsTests { #expect(refused != nil) } } + +/** + A conversation that succeeds, against a real model — a local Ollama, so no key is needed. + + Skipped unless `XBOT_LIVE_OLLAMA_MODEL` is set as well, to a pulled model that can call tools + (`qwen2.5:7b` works). Everything above proves how a turn fails; this is the one place a turn is + proven to answer, remember within a conversation, drive its computer through the client-tool loop, + and leave the thread holding each message once. Launch checklist item 5, the engine half. + + A model is not deterministic, so the prompts ask for things a small one reliably gets right and the + assertions check for a word, never a sentence. + */ +@Suite(.enabled(if: ProcessInfo.processInfo.environment["XBOT_LIVE_ENGINE_URL"] != nil + && ProcessInfo.processInfo.environment["XBOT_LIVE_OLLAMA_MODEL"] != nil)) +struct LiveConversationTests { + private let env = ProcessInfo.processInfo.environment + + private var client: HTTPEngineClient { + let url = URL(string: env["XBOT_LIVE_ENGINE_URL"] ?? "http://127.0.0.1:3001")! + let token = env["XBOT_LIVE_ENGINE_TOKEN_FILE"].flatMap { + try? String(contentsOfFile: $0, encoding: .utf8).trimmingCharacters(in: .whitespacesAndNewlines) + } + return HTTPEngineClient(baseURL: url, token: token) + } + + private struct Turn { + var text = "" + var tools: [String] = [] + var failure: String? + } + + private func turn(_ text: String, in channel: Channel.ID) async throws -> Turn { + var turn = Turn() + for try await event in client.send(text, to: channel) { + switch event { + case .textDelta(_, let delta): turn.text += delta + case .toolCall(_, _, let name, _): turn.tools.append(name) + case .failed(_, let reason): turn.failure = reason + default: continue + } + } + print("live turn:", text, "→", turn.tools, turn.failure ?? turn.text) + return turn + } + + @Test func aConversationAnswersRemembersUsesItsComputerAndIsKeptOnce() async throws { + let client = client + // What `ModelProviderCatalog.engineRouting(for: "ollama")` sends, which lives in XBotCore. + let agent = try await client.createAgent(AgentDraft( + name: "Conversation check", + model: ModelSelection( + provider: "Ollama", providerID: "openai-compatible", model: env["XBOT_LIVE_OLLAMA_MODEL"]!, + baseURL: env["XBOT_LIVE_OLLAMA_BASE_URL"] ?? "http://host.docker.internal:11434/v1", capabilities: [] + ) + )) + let channel = try await client.createChannel(agentIds: [agent.id]) + + let told = try await turn("Remember this codeword: tangerine. Reply with just OK.", in: channel.id) + #expect(told.failure == nil) + #expect(!told.text.isEmpty) + + // A 7B model sometimes writes the call out as text instead of making it. That is the model's + // choice, not the loop failing, so it gets one more ask before the loop is judged. + var sent = 2 + var browsed = try await turn( + "Use your browser to open https://example.com, then tell me the page's title.", in: channel.id + ) + if !browsed.tools.contains("computer_navigate") { + sent += 1 + browsed = try await turn("Call the computer_navigate tool with https://example.com now.", in: channel.id) + } + #expect(browsed.failure == nil) + #expect(browsed.tools.contains("computer_navigate")) + #expect(browsed.text.localizedCaseInsensitiveContains("example")) + + // Two turns and a tool loop later. Each run carries the whole conversation, so this is memory. + let asked = try await turn("What was the codeword I gave you? Answer with the word only.", in: channel.id) + #expect(asked.failure == nil) + #expect(asked.text.localizedCaseInsensitiveContains("tangerine")) + + // The tool loop resends messages the thread already holds and relies on their ids matching. + let history = try await client.messages(in: channel.id) + print("live history:", history.map { "\($0.isFromUser ? "user" : "agent") \($0.id) \($0.text.prefix(40))" }) + #expect(Set(history.map(\.id)).count == history.count) + #expect(history.filter(\.isFromUser).count == sent + 1) + } +} diff --git a/docs/12-roadmap.md b/docs/12-roadmap.md index cae308c..0ca6108 100644 --- a/docs/12-roadmap.md +++ b/docs/12-roadmap.md @@ -141,7 +141,9 @@ and `openai-compatible` — rather than two vendors. An end-to-end conversation through the server no longer needs Intelligence credentials: with none of the `INTELLIGENCE_*` variables set, `copilot.ts` runs on `LocalThreadRunner` (ADR-0008). That run -through the server is launch checklist item 5, and has not been done yet. +through the server is launch checklist item 5. Its engine half was run on 27 September against a +local Ollama model, and found three faults on this hop that every conversation hit; all three are +fixed and recorded there. **Still open:** a second live vendor. The three native adapters are built and each is covered by a test asserting which client a selection gets (`agent-langgraph/tests/models-build.test.ts`). Usage accounting is done — the agent sums `usage_metadata` across a turn and emits diff --git a/docs/13-launch-checklist.md b/docs/13-launch-checklist.md index 66e3885..c69a60c 100644 --- a/docs/13-launch-checklist.md +++ b/docs/13-launch-checklist.md @@ -134,11 +134,33 @@ engine (`docs/plans/computer-client-tools.md`): **There is an automated version of the engine half.** `apps/mac/Tests/XBotEngineTests/LiveEngineTests.swift` drives the real client against a running engine — point it at a throwaway one, since it creates -agents. Its conversation-store assertions now run against local mode by default; a separate -`XBOT_LIVE_ENGINE_HAS_INTELLIGENCE=1` branch exists only for a deployment that configures -Intelligence by hand, which nothing here needs. This suite, and `LocalThreadRunner`'s own -integration test (`engine/server/tests/local-thread-runner.integration.test.ts`), still need to be -run before this item can be checked off — see the verification commands in `CLAUDE.md`. +agents. **It has now been run** (27 September), against an engine built from the working tree and a +local Ollama (`qwen2.5:7b`): set `XBOT_LIVE_OLLAMA_MODEL` and `LiveConversationTests` holds a real +three-turn conversation — it answers, drives `computer_navigate` through the client-tool loop and +reports the page, remembers a codeword two turns later, and leaves each message in the thread once. +Restarting that engine left every conversation's message count unchanged. What remains for a person +is the app half: the same conversation typed into the window, and the asks. + +That first run found three engine faults, each of which failed every conversation: + +- **The managed Bot could not be dialled.** `createAgentFetch` refused `127.0.0.1:4201`, the + address the image itself gives the Bot, so every run through the server ended in "This deployment + will not dial…". The Bot's configured host is now added to the allowed hosts + (`server/src/agents/managed-agent-host.ts`), and only that host and port. +- **A keyless `openai-compatible` endpoint was refused locally.** OpenAI's client falls back to + `OPENAI_API_KEY` when handed no key, so every Ollama run failed with "Missing credentials". It now + gets a placeholder the endpoint ignores (`agent-langgraph/src/models/build.ts`). +- **Every Anthropic run was refused.** The server's standing role message arrived after the Bot's + guidance, and Anthropic accepts a system message only first. All system texts are now merged, in + order, into one leading message (`agent-langgraph/src/history.ts`). + +**One open finding.** On a freshly started engine, roughly one full live run in six had a turn come +back empty, or a request go unanswered, while all nine live tests ran at once. The engine logged +"Thread already running" or, once, "No agents are registered", both of which should be impossible +there. It was never seen with the conversation suite alone (5 of 5 fresh engines), nor with a +logging proxy in the path (12 of 12). A leftover process, a database lock, identity, idle +keep-alive and split request writes were each checked and ruled out. Worth one more look before +launch, starting from the app under ordinary use rather than nine tests at once. **Also close the last M2 item here:** point one agent at a second real vendor — Anthropic is already proven, so use OpenAI or Google — and confirm the reply comes from the vendor you picked. A model @@ -197,6 +219,10 @@ From a shell holding no credentials: passed all eight live client tests: health, read endpoints, the vault, agent and conversation round-trip, browser, files, shell, and the asks. That run predates ADR-0008; the conversation store now needs re-running against `LocalThreadRunner` rather than a CopilotKit key — see item 5. +- On 27 September, it was: a throwaway engine in local mode with no CopilotKit key held a real + conversation with a local Ollama model, used its computer through the client-tool loop, remembered + across turns, and kept every conversation intact through a restart. Item 5 has the details and + the one open finding. - The pinned engine is one multi-arch image (linux/amd64 and linux/arm64). On 14 September an Apple Silicon Mac pulled it anonymously, got arm64, was healthy in about ten seconds on a port other than 3001, and passed the live suite: browser, files, shell, and a help request handed back. Before that diff --git a/docs/plans/computer-client-tools.md b/docs/plans/computer-client-tools.md index 07a6a93..1c292dd 100644 --- a/docs/plans/computer-client-tools.md +++ b/docs/plans/computer-client-tools.md @@ -32,11 +32,14 @@ been offered a browser, files, or a shell — the product's headline feature doe - **C1 done**, verified against a real engine: navigate, snapshot, write/read a file, run a command (`LiveComputerToolsTests`). - **C2 done.** Every run carries `ComputerTools.wireTools`. -- **C3 done, verified only against a stubbed engine.** `WireTranscript` rebuilds the conversation with +- **C3 done, verified against a real engine** on 27 September (`LiveConversationTests`, local + Ollama, no CopilotKit key): the follow-up run continues after `computer_navigate`, and the thread + holds each message once. + `WireTranscript` rebuilds the conversation with `@ag-ui/client`'s reducer, id for id, starting from the thread's stored history — which also fixes agents having no memory of earlier messages. `HTTPEngineClient.stream` runs pending calls at - `RUN_FINISHED` and follows up, capped at 25 rounds. **Needs one real conversation with a CopilotKit - key** before v1: that the history route's rows round-trip, and that follow-ups do not duplicate. + `RUN_FINISHED` and follows up, capped at 25 rounds. The real conversation this waited on no longer + needs a CopilotKit key (ADR-0008), and has been held: rows round-trip and follow-ups do not duplicate. - **C4 done, verified against a stubbed engine.** While a turn runs, `AppState` polls `/control` every two seconds, as upstream's client does. A help request shows "needs you" with Take control, which opens the screen; handing back ends the wait. A secret request shows a masked field that posts diff --git a/engine/agent-langgraph/src/history.ts b/engine/agent-langgraph/src/history.ts index 2f98252..75a5132 100644 --- a/engine/agent-langgraph/src/history.ts +++ b/engine/agent-langgraph/src/history.ts @@ -24,6 +24,14 @@ export { NO_ANSWER_CAME }; /** Translate the conversation AG-UI carries into LangChain's message classes. */ export function toLangChainMessages(input: RunAgentInput): BaseMessage[] { const messages: BaseMessage[] = [new SystemMessage(COMPUTER_GUIDANCE)]; + /* + * Every system text, gathered into that first message rather than pushed where it arrived. + * + * The server sends standing messages of its own — the role, what this Bot holds. OpenAI answers + * system messages anywhere; Anthropic refuses the whole run: "System messages are only permitted + * as the first passed message". Merged in order, so what each one says is unchanged. + */ + const system = [COMPUTER_GUIDANCE]; /* * Which calls in this history were ever answered. @@ -53,7 +61,7 @@ export function toLangChainMessages(input: RunAgentInput): BaseMessage[] { continue; } if (message.role === "system" || message.role === "developer") { - messages.push(new SystemMessage(String(message.content ?? ""))); + system.push(String(message.content ?? "")); continue; } if (message.role === "tool") { @@ -123,6 +131,7 @@ export function toLangChainMessages(input: RunAgentInput): BaseMessage[] { messages.push(new HumanMessage(CONTINUE_TURN)); } + messages[0] = new SystemMessage(system.join("\n\n")); return messages; } diff --git a/engine/agent-langgraph/src/models/build.ts b/engine/agent-langgraph/src/models/build.ts index 23be3fd..510e59a 100644 --- a/engine/agent-langgraph/src/models/build.ts +++ b/engine/agent-langgraph/src/models/build.ts @@ -80,7 +80,11 @@ export function buildChatModel({ */ return new ChatOpenAI({ model: resolved.model, - apiKey: resolved.apiKey, + // A keyless compatible endpoint (a local Ollama) still needs *a* key: without one the client + // falls back to OPENAI_API_KEY and refuses the run. The endpoint ignores whatever is sent. + apiKey: + resolved.apiKey || + (resolved.providerId === "openai-compatible" ? "not-needed" : undefined), streaming: true, ...(resolved.baseURL ? { configuration: { baseURL: resolved.baseURL } } diff --git a/engine/agent-langgraph/tests/history-system.test.ts b/engine/agent-langgraph/tests/history-system.test.ts new file mode 100644 index 0000000..b56bc6c --- /dev/null +++ b/engine/agent-langgraph/tests/history-system.test.ts @@ -0,0 +1,38 @@ +import { describe, expect, test } from "bun:test"; +import type { RunAgentInput } from "@ag-ui/core"; +import { SystemMessage } from "@langchain/core/messages"; +import { COMPUTER_GUIDANCE } from "../../shared/bot-prompt"; +import { toLangChainMessages } from "../src/history"; + +/** + * Every system text arrives as one leading system message. + * + * The server sends its own standing messages — the role, what this Bot holds — and the Bot puts its + * guidance first. OpenAI answers any number of system messages anywhere; Anthropic refuses the run + * outright ("System messages are only permitted as the first passed message"), so every Anthropic + * agent failed before reaching the vendor. Found by the first live run through the server. + */ +const input = (messages: unknown[]): RunAgentInput => + ({ messages }) as unknown as RunAgentInput; + +describe("system messages", () => { + test("are merged, in order, into the first message and nowhere else", () => { + const messages = toLangChainMessages( + input([ + { role: "system", content: "You are the research agent." }, + { role: "user", content: "hello" }, + { role: "developer", content: "You hold Drive." }, + ]), + ); + + const system = messages.filter((m) => m instanceof SystemMessage); + expect(system).toHaveLength(1); + expect(messages[0]).toBe(system[0]); + const text = String(system[0]?.content); + expect(text.indexOf(COMPUTER_GUIDANCE)).toBe(0); + expect(text.indexOf("You are the research agent.")).toBeLessThan( + text.indexOf("You hold Drive."), + ); + expect(messages).toHaveLength(2); + }); +}); diff --git a/engine/agent-langgraph/tests/models-build.test.ts b/engine/agent-langgraph/tests/models-build.test.ts index d5b1757..ef427da 100644 --- a/engine/agent-langgraph/tests/models-build.test.ts +++ b/engine/agent-langgraph/tests/models-build.test.ts @@ -54,6 +54,21 @@ describe("the client a selection is answered by", () => { ); }); + test("a keyless openai-compatible endpoint still hands the client a key", () => { + // A local Ollama asks for none, but OpenAI's client falls back to OPENAI_API_KEY and refuses the + // run without one: "Missing credentials", on every turn, found by the first live Ollama run. + const model = buildChatModel({ + selection: { + providerId: "openai-compatible", + model: "qwen2.5:7b", + baseURL: "http://host.docker.internal:11434/v1", + }, + fallback: undefined, + keys: {}, + }) as ChatOpenAI; + expect(model.apiKey).toBeTruthy(); + }); + test("an agent that never chose inherits the workspace default", () => { const model = buildChatModel({ selection: undefined, diff --git a/engine/server/src/agents/managed-agent-host.ts b/engine/server/src/agents/managed-agent-host.ts new file mode 100644 index 0000000..d19f119 --- /dev/null +++ b/engine/server/src/agents/managed-agent-host.ts @@ -0,0 +1,20 @@ +/** + * The addresses every run may dial: the ones the deployment named, plus its own Bot's. + * + * The one-container image runs the Bot beside the API at a loopback address, which the + * private-address guard in `endpoint.ts` refuses — correctly, for an address a person typed. This + * one came from the deployment's own configuration (`MANAGED_AGENT_AG_UI_URL`), which is exactly + * what `allowedHosts` exists for. Host and port together, so the rest of loopback — the database, + * the computer — stays refused. Compose never needed this: there the Bot is a service name, which is + * not a private literal. + * + * Dialling only. Registration keeps the named list alone, so nobody can register an agent at the + * managed Bot's address by way of this. + */ +export function dialableHosts( + named: ReadonlySet, + managedAgent: { endpoint: URL } | undefined, +): ReadonlySet { + if (!managedAgent) return named; + return new Set([...named, managedAgent.endpoint.host]); +} diff --git a/engine/server/src/index.ts b/engine/server/src/index.ts index d7f31d4..eb6af4c 100644 --- a/engine/server/src/index.ts +++ b/engine/server/src/index.ts @@ -8,6 +8,7 @@ import { eq } from "drizzle-orm"; import { COMPUTER_GUIDANCE } from "../../shared/bot-prompt"; import { mintRunAssertion, readRunAssertion } from "./agents/callback-token"; import { createAgentFetch } from "./agents/endpoint"; +import { dialableHosts } from "./agents/managed-agent-host"; import { askTheirOwnPerson, escalationTool } from "./agents/escalation"; import { createHandoffDesk, HANDOFF_KIND } from "./agents/handoff"; import { createHandoffDelivery } from "./agents/handoff-delivery"; @@ -603,8 +604,12 @@ const selectionForActor = (actorId: string): ToolSelection => ({ // reading and the same one `createApp` takes. const agentFetch = createAgentFetch({ allowPrivateHosts: config.computer?.allowPrivateHosts === true, - // Named addresses are reachable on every hop, not only the one that was registered. - allowedHosts: config.agentEndpointAllowedHosts, + // Named addresses are reachable on every hop, not only the one that was registered. The + // deployment's own Bot is one of them: in the one-container image it sits on loopback. + allowedHosts: dialableHosts( + config.agentEndpointAllowedHosts, + config.managedAgent, + ), // The refusal is what the run already knows; this is what the deployment knows. Written here // rather than in `endpoint.ts` so that file keeps deciding and nothing else, the way the // target check it reuses does. diff --git a/engine/server/tests/managed-agent-dial.test.ts b/engine/server/tests/managed-agent-dial.test.ts new file mode 100644 index 0000000..b9613ea --- /dev/null +++ b/engine/server/tests/managed-agent-dial.test.ts @@ -0,0 +1,40 @@ +import { describe, expect, test } from "bun:test"; +import { createAgentFetch } from "../src/agents/endpoint"; +import { dialableHosts } from "../src/agents/managed-agent-host"; + +// The one-container image runs the Bot beside the API at a loopback address (docker/s6 api/run). +// Every run dials it through the private-address guard, which refused it: no conversation through +// the server could ever be answered. +const managed = { + endpoint: new URL("http://127.0.0.1:4201/ag-ui"), + token: "t", +}; +const answering = (async () => + new Response("ok", { status: 200 })) as unknown as typeof fetch; + +describe("the deployment's own Bot", () => { + test("is dialled", async () => { + const dial = createAgentFetch({ + allowedHosts: dialableHosts(new Set(), managed), + fetchImpl: answering, + }); + expect((await dial("http://127.0.0.1:4201/ag-ui")).status).toBe(200); + }); + + test("names its port, so the rest of loopback stays refused", async () => { + const dial = createAgentFetch({ + allowedHosts: dialableHosts(new Set(), managed), + fetchImpl: answering, + }); + await expect(dial("http://127.0.0.1:5432/")).rejects.toThrow(); + }); + + test("adds nothing when there is none, and keeps what was named", () => { + const named = new Set(["agents.internal"]); + expect(dialableHosts(named, undefined)).toBe(named); + expect([...dialableHosts(named, managed)].sort()).toEqual([ + "127.0.0.1:4201", + "agents.internal", + ]); + }); +});