A remote control plane for every machine you own: install Browserland on each one, and a single browser tab becomes a desktop of live terminals across all of them — including the AI coding agents running inside.
Point a browser at one broker and you get a full windowed desktop of live terminals: tile them, tab them, split them, and drag them across virtual desktops. Each terminal is a real PTY running on some machine, streamed to the browser over a WebSocket. Add your other machines as hosts (token-auth'd, e.g. over Tailscale) and their terminals appear in the same desktop — launch a shell on the laptop, the desktop, and the server from one + menu, side by side.
The shells keep running even when no browser is attached. Close the tab, come back tomorrow, and the screen heals from a snapshot — exactly where you left it.
The name says it plainly: a whole little desktop — windows, terminals, and your
fleet of AI coding agents — living entirely in a browser tab. (webterm is the
Python package and module name.)
Browserland also exposes an MCP server, letting LLM harnesses drive the terminals directly — including full-screen TUIs.
Browserland used to treat loopback as an auth exemption — with no
auth_tokenset (the shipped default), anything on the machine could read every terminal's title/pid/cwd, attach to any PTY, and reach host-wide file read/write. That was never sound:tailscale servein front of a127.0.0.1bind makes every tailnet request arrive from loopback, and any web page you have open can reach loopback too. A token is now required on every route and every interface, with no exemption and no opt-out.New install? Nothing to do. The broker mints its own token on first run and prints a ready-to-open URL. You never have to create one.
Were you running without a token? Then on your next start:
- The browser will ask for a password. Get it with
python -m webterm.broker --print-token— it's also minted intowebterm_token.jsonbeside your state store.- Terminals launched by the previous run cannot reconnect and must be relaunched once. Their shells keep running, but they were started without the token and it can't be injected into a live process. This is the one unavoidable loss — finish or accept whatever is in them first.
- Scripts calling
/sessions,/launchor/file/*over loopback now get401. SendAuthorization: Bearer <token>.Running a fleet? Set
auth_tokenyourself inbroker_config.jsonbefore the cutover, so every host shares one password you already know instead of a different minted value per machine.Full details, a does-this-affect-me table, and recovery steps:
wiki/Upgrading.md.
Browserland is two small programs and a browser:
- Agents (producers) are headless processes that own a real terminal —
pty.openptyon Linux, ConPTY/WinPTY on Windows. An agent runs a command (bash,cmd.exe, an AI coding agent, anything), keeps a ring buffer of recent output, and streams the terminal over a single WebSocket. - The broker is a small web server. Agents register with it; browsers connect to it. It relays bytes both ways, serves the desktop UI, and can spawn new agents on demand from a list of pre-approved profiles.
- The browser renders the desktop — a tiling window manager over xterm.js — and sends your keystrokes back to the PTY.
The wire format is a deliberately small set of JSON frames over one WebSocket, so any producer that speaks them can register with the broker — Browserland's own agent is just the reference implementation.
┌─────────┐ binary ANSI + JSON ┌────────┐ /ws?session=<id> ┌─────────┐
│ agent │ ──────────────────▶ │ broker │ ──────────────────▶ │ browser │
│ (PTY + │ ◀────────────────── │ │ ◀────────────────── │ xterm.js│
│ ConPTY/ │ input/resize/ └────────┘ input/paste/resize └─────────┘
│ openpty)│ snapshot_please
└─────────┘
Because the PTY lives in the agent, not the browser, terminals survive browser reloads and even broker restarts — the agent reconnects with backoff and the browser's attach triggers a snapshot redraw from the ring buffer.
The part people miss: Browserland is a remote control plane, not just a terminal for the machine it runs on. Run a broker on each of your machines, then register them as hosts in the UI (Control Panel → Hosts — a name, a URL, and that broker's token; a Tailscale address works great). One browser tab then fronts the whole fleet:
one browser tab
│
┌──────────────────┼──────────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌─────────┐
│ broker │ │ broker │ │ broker │ one per machine
│ laptop │ │ desktop │ │ server │ (LAN / Tailscale)
└─┬─────┬─┘ └─┬─────┬─┘ └─┬─────┬─┘
▼ ▼ ▼ ▼ ▼ ▼
shell claude shell build shell claude your PTYs + agents
- The right-click + menu groups launch profiles per host — open a PowerShell on the Windows box and a zsh on the Linux server from the same menu, and tile them next to each other.
- Every window is badged with its host; the taskbar's broker status — a chip for one host, a single aggregate badge for several — shows at a glance which brokers are reachable, with per-broker detail in the + menu.
- Each host keeps its own token, profiles, and file API — the UI fans out, the brokers stay independent. A machine going down takes its windows stale, not the desktop.
- The MCP server speaks the same multi-host language: one
--hostslist lets an AI harness enumerate and drive terminals across all brokers through a single tool surface.
wiki/Setup-and-Onboarding.md walks through joining machines step by
step.
The desktop with tiled and floating terminals
- In-browser tiling window manager (niri-style): floating and tiled windows, a virtual-desktop pager, and a taskbar — all in the page, no native app.
- Nested tabs and splits: any tile can hold a tab group, and tabs can hold splits, via a recursive cell model. Build the layout you want by dragging.
- More than terminals: sticky notes, a CodeMirror 6
text editor, a file manager, a task manager, a synced scratchpad, a clipboard
history, and a session recorder — each a bundled mod (see
Mods) over the broker's token-gated
/file/*and per-mod store APIs. - Session recording & replay: hit ⏺ on any terminal to record it byte-faithfully, then play it back at the original size — pause, 0.25×–8× speed, continuous rewind, and notes pinned to timestamps.
- A bundled mod system: the desktop's app windows and widgets are mods
over a small, versioned
ctxAPI — a mod is one folder with a manifest and an entry script, and it can add window kinds, per-terminal title-bar widgets, taskbar chips, Control-Panel settings, and in-app help pages. - Cross-platform PTY: Linux
pty.openpty; Windows auto-selects ConPTY or WinPTY — even headless/detached agents acquire a hidden console and re-enable Ctrl-C so they run ConPTY with a working interrupt and live resize, falling back to WinPTY only if console acquisition fails (or when forced via--pty-backend winpty). - AI agent fleet: detects the foreground coding agent in each window
(
claude/codex/grok/opencode/hermes) and tracks live OSC title + working directory, with an opt-in per-window git-status widget (thegitmod). - Multi-host — the control plane: attach the same UI to the brokers on all your other machines (e.g. over Tailscale) and run their terminals side by side in one tab, with live broker status in the taskbar and launch menu. See One tab, every machine.
- Single active-browser lease: exactly one browser drives input at a time, so two open tabs never fight over the keyboard.
- Opt-in MCP / AI agent access: an MCP client or AI harness can list, observe, drive, and launch terminals — including live interactive TUIs and the console you're working in — under per-window access modes you control. See MCP & AI agent access below.
- Token auth, no open RCE: one token gates every connection — every route, every interface, loopback included — and doubles as the UI password. Nothing to set up: a broker with no token configured mints one and prints the URL to open. Launching is profiles-only (the client can never supply a raw command).
- Launch profiles you can edit in the UI: add WSL/zsh/PowerShell shells from Control Panel → Launch profiles (one-click WSL-distro / shell Detect…), applied live with no restart — still profiles-only. See wiki/Launch-Profiles.md for recipes and the security model.
The desktop's app windows and most of its optional chrome ship as mods —
the terminal pipeline, window manager, the multi-host connection model, and
MCP surfaces stay core (only the optional sharing of that host list across
browsers is a mod). A mod is a self-contained folder under webterm/broker/mods/<id>/
holding a manifest (mod.json), its script(s), and optional CSS + an in-app
help page. The bundled ones are trusted first-party code, spliced from an
explicit allow-list into the single served page; a broker can also be handed a
mod at runtime (POST /mods/install, ids prefixed x-), which is served from
its own same-origin URL and picked up on the next page load. Either way a mod
registers through a versioned ctx API that exposes per-terminal-window hooks,
app-window kinds (which appear in the + launch menu), taskbar chips,
Control-Panel settings, a durable per-mod server store with revision history, an
in-desktop copy/paste observer, and help cards.
A mod is not sandboxed — it is same-origin code with the page's full
authority, and ctx is a set of reviewed choke points rather than a boundary.
Writing one (and the trust model behind installing one) is covered in
wiki/Writing-a-Mod.md.
A mod's settings sync across your browsers via the broker's shared state;
its enable/disable toggle is deliberately per-browser — flip it in
Control Panel → Mods. A broker-side mods_enabled master switch gates
the whole system, and a broker can pin individual mods on or off for every
browser that loads its page. The desktop ships with nineteen:
| Mod | What it adds | Default |
|---|---|---|
theme |
color-scheme picker for the desktop chrome | on |
pattern |
desktop background patterns | on |
clock |
taskbar clock with timezone picker | on |
help |
the in-app help window + ? chip | on |
sticky |
sticky notes | on |
editor |
CodeMirror 6 text editor | on |
file-manager |
dual-pane file manager | on |
task-manager |
live per-host process list | on |
scratchpad |
server-backed notes, synced across browsers with revision history | on |
recorder |
terminal session recorder + fixed-size player (speed, continuous rewind, timestamped notes) | on |
host-registry |
optional shared broker list — publish/pull your host list across browsers | on |
mousemode |
title-bar chip while a full-screen app owns the mouse | on |
mod-sync |
copy this broker's mod setup to your other brokers (all or selected) | on |
workspaces |
vertical workspaces — taskbar pager, per-workspace names, Send to workspace | on |
git |
per-terminal git branch + dirty-state widget | off |
aistatus |
AI-provider status chip + window | off |
clipboard |
rolling history of copies/pastes made through the desktop | off |
termfont |
terminal font picker | off |
clipboard and aistatus are off by default on principle — clipboards carry
secrets, and status polling talks to third-party endpoints; git and
termfont are simply opt-in preferences. Enabling any of them is one click
in the Mods pane.
The broker exposes a token-gated /mcp/* HTTP API and ships a stdio
MCP server (webterm.mcptool), so any MCP
client or AI harness can list, observe, drive, and launch terminals. The
agents are just producers; the broker stays the sole authority — every MCP call
is gated by the same per-window access modes and the master enable switch.
- Interactive TUIs, as plain text —
read_screenrenders the current screen of a live terminal by replaying its PTY ring buffer through pyte, so a harness can read full-screen apps (btop, htop, vim, less) — not just line-oriented scrollback — andsend_inputtypes into them. Without pyte the read falls back to a dependency-free in-house renderer that still returns a bounded rendered grid (degraded: trueis now reserved for a rare last-ditch raw decode). - The live session you're working in —
list_terminalsenumerates running sessions (id, title, cwd, agent, kind, cols/rows, mode), so a harness can attach to the exact console a person is using right now, read its state, and (inreadwrite) drive it. Sessions persist across browser reloads and broker restarts, so the handle stays valid.
Each tool maps to a broker endpoint and returns its JSON. Window ids are
namespaced "<host>:<int>" strings so one server can front several brokers (see
Multi-host below); with a single broker the host is default
("default:12345").
| Tool | Endpoint | Notes |
|---|---|---|
mcp_info(host?) |
GET /mcp/info |
feature flags (allow_launch, default_mode). Omit host → dict keyed by host name |
list_terminals |
GET /mcp/terminals |
{"terminals":[…], "errors":{host:msg}}: all hosts merged, each terminal's host set to the config name (broker's machine hostname preserved as machine_host) + namespaced id; a down host lands in errors without sinking the rest. Each terminal carries a pyte flag — false = the agent lacks pyte, so attr_runs/keyframe-repair are unavailable and sparse alt-screen frames are flagged partial only (#134) |
list_profiles(host?) |
GET /mcp/profiles |
launchable profile names + default. Omit host → dict keyed by host name |
read_screen(id) |
POST /mcp/read |
screen rendered as a bounded plain-text grid (pyte, or a dependency-free fallback); result carries content_hash, stable_hash (the cursor-blind digest — a cursor blink in place doesn't change it, a cursor move does), and idle_ms (best-effort ms since the last PTY output — absent from older agents, unreliable for a perpetually-animating app). Waits (one call, timeout_ms-bounded, mutually exclusive): wait_for_change / wait_for_text / wait_for_regex / wait_for_idle (block until the screen settles — stable_hash unchanged for N ms) |
send_input(id, data) |
POST /mcp/input |
target window must be in readwrite mode; newlines are sent as Enter (CR) so commands run (incl. on PowerShell) |
send_keys(id, keys, delay_ms?) |
POST /mcp/input |
send control/escape keys — ["C-c"], ["Esc"], ["Up","Enter"] — that plain text can't express; delay_ms (or a per-terminal set_pace default) writes one key per POST for a frame-polling TUI |
set_pace(id, pace_ms) |
POST /mcp/pace |
readwrite; set a per-terminal DEFAULT send_keys pacing (ms, capped 1000, 0 disables) so multi-key sends auto-pace without passing delay_ms — for a frame-polling TUI (Dwarf Fortress). Broker-local + ephemeral (resets on agent reconnect) |
reset_terminal(id) |
POST /mcp/reset |
readwrite; correlated round-trip that wipes the agent's screen-render buffer so the next read_screen starts clean (502 on a non-agent producer) |
flush_input(id) |
POST /mcp/flush |
readwrite; correlated round-trip that discards keystrokes queued to the app but not yet consumed — the input-side mirror of reset (502 on a non-agent producer; a no-op on a Windows/ConPTY agent) |
launch_terminal(profile?, cols=80, rows=24, title?, cwd?, host?) |
POST /mcp/launch |
broker must have allow_launch enabled; host required when multiple hosts are configured |
Multi-host. Pass --hosts (or $BROWSERLAND_MCP_HOSTS) a JSON array of
{name, url, token} descriptors to serve N brokers from one server process;
every id-taking tool routes on the "<host>:…" prefix. The single
--broker-url/--token form is the one-host shorthand (default). See
webterm/mcptool/README.md for details.
Scopes. Several MCP clients can share one broker without seeing each other's
windows: a server started with --scope projA (or BROWSERLAND_MCP_SCOPE, or a
per-host "scope") sees only the windows tagged projA, while a server with no
scope keeps the admin view of everything. A window that server launches is
tagged for it and starts in readwrite; tag a hand-started one from its robot
button or title-bar menu. Control Panel → Access → MCP access → Copy
.mcp.json writes a ready-to-save client file for a scope. A scope is a
convention, not security (every client holds the same token), and a window's
scope and MCP mode now survive a broker restart. See
wiki/MCP-and-AI-Agents.md.
The
send_inputtool maps newlines indatato a carriage return — the byte a real Enter key sends — so a line submits on PowerShell/PSReadLine (which treats a bare line-feed as a soft continuation) and on a Unix shell alike. The rawPOST /mcp/inputendpoint stays verbatim: drive it directly to send a literal LF or hand-crafted control/escape bytes.
Access is layered and opt-in — nothing is reachable until you turn it on:
- Master enable is off by default; while off, every
/mcp/*call returns403 mcp_disabled. - Per-window mode is
off/read/readwrite, with a globaldefault_modefor new windows (a scopedlaunch_terminalstarts inreadwriteunless it passes amode, and a mode set on a window survives broker restarts).offhides a window entirely;readallows observation;readwriteadditionally allowssend_input. allow_launchis a separate gate forlaunch_terminal.- The MCP token is a bearer secret distinct from the browser
auth_token(the UI password): the broker pins it withWEB_TERMINAL_MCP_TOKEN(or thewebterm_mcp.jsonsidecar), and the MCP client passes the same secret viaBROWSERLAND_MCP_TOKEN(as the harness examples below show).
claude mcp add browserland \
--env BROWSERLAND_MCP_TOKEN=… \
--env BROWSERLAND_MCP_URL=http://127.0.0.1:4445 \
-- python -m webterm.mcptoolAny other stdio MCP client (Hermes, your own, …) registers the same way — point it at the launch command and pass the two env vars:
{
"mcpServers": {
"browserland": {
"command": "python",
"args": ["-m", "webterm.mcptool"],
"env": {
"BROWSERLAND_MCP_TOKEN": "…",
"BROWSERLAND_MCP_URL": "http://127.0.0.1:4445"
}
}
}
}Or run the server directly, talking to the local broker:
BROWSERLAND_MCP_TOKEN=… python -m webterm.mcptool
BROWSERLAND_MCP_TOKEN=… browserland-mcp --broker-url http://127.0.0.1:4445For the full HTTP contract, error table, and config sidecar, see
wiki/Technical-Reference.md and
webterm/mcptool/README.md.
You need Python ≥ 3.9. Install from a checkout:
pip install -e .# broker (default 127.0.0.1:4445) — mints a token on first run and prints
# the URL to open, token included
python -m webterm.broker
# the token again, any time (never mints one):
python -m webterm.broker --print-token
# an agent running cmd.exe, registered with the local broker
python -m webterm.agent -- cmd.exe
# then open the printed http://127.0.0.1:4445/?token=... and click the
# session (or "new terminal")Windows agents also need a PTY backend: pip install -e ".[windows]" (pulls in
pywinpty).
./launchers/run-broker.sh # broker
./launchers/run-agent.sh -- bash -l # agentThe launcher scripts bootstrap a virtualenv with the runtime dependencies on
first run. The broker prints the URL to open, token included; python -m webterm.broker --print-token reprints it later. run-agent.sh picks the token
up from webterm_token.json automatically, so a hand-started agent on the same
box needs no setup.
pip install -e . # core: broker + agent
pip install -e ".[windows]" # + pywinpty (Windows PTY backend)
pip install -e ".[pyte]" # + pyte (tier-2 snapshot rendering)
pip install -e ".[procs]" # + psutil (task manager, agent badge, live cwd)
pip install -e ".[mcp]" # + the stdio MCP server (Python ≥ 3.10)
pip install -e ".[dev]" # + pytest for the test suitepsutil (the procs extra) is best-effort: it powers the task-manager
process list, the foreground-agent badge, and live-cwd tracking.
Without it the agent still runs and still destroys windows — those three views
just degrade (empty list / no badge / no cwd).
Mix and match, e.g. pip install -e ".[pyte,mcp,dev]". The mcp extra requires
Python ≥ 3.10 (the MCP SDK), while everything else runs on Python ≥ 3.9.
See MCP & AI agent access for running the server.
| Path | What |
|---|---|
webterm/agent/ |
Headless producer: PTY backends, output ring buffer, OSC-title sniffer, reconnecting WebSocket client |
webterm/broker/ |
Web server: desktop UI (ui.py assembles the served page from ordered *.html/*.css/*.js fragments), /ws relay, producer WS, session list, profiles-only launch |
webterm/broker/mods/ |
The bundled desktop mods — one folder per mod: mod.json manifest, script(s), optional CSS + in-app help page (authoring guide: wiki/Writing-a-Mod.md) |
webterm/mcptool/ |
The shipped stdio MCP server wrapping the broker's /mcp/* API |
webterm/protocol.py |
The single source of truth for the JSON frame shapes |
launchers/ |
venv-bootstrapping run scripts (and systemd units) for both OSes |
tests/ |
pytest suite: protocol, snapshots, agent↔broker integration, real-PTY round trips |
Everything lives in wiki/ — one tree, and the same source the
in-app Help window reads. Open Help with the ? chip in the taskbar to search
it without leaving the desktop; the developer and operator pages are in there
too, behind Include developer docs.
| Page | |
|---|---|
| Start here | Home · Getting Started · Keyboard Shortcuts |
| Building layouts | Window Modes · Arranging Windows · Columns & Widths · Snapping & Pop-out · Floating Window Controls |
| The desktop | Workspaces · Taskbar · Context Menus · Window Types |
| Multi-host & AI | Hosts & Multi-Browser · MCP & AI Agents |
| For developers | Setup & Onboarding · Launch Profiles · Writing a Mod · Technical Reference · Upgrading |
New here, or setting up more than one machine? Start with
wiki/Setup-and-Onboarding.md — the onboarding guide: the broker / agent /
browser mental model, joining machines over Tailscale via Control Panel → Hosts,
what not to hand-edit, and running the broker unattended in the background
(Windows Task Scheduler / Linux systemd). It's written to be followed by a human
or a coding agent.
Upgrading an existing install? wiki/Upgrading.md is
the breaking-change record — what changed, whether it affects you, and how to
recover. Most recently: a token is now required on every connection,
including loopback (#142), which strands terminals launched by a previously
tokenless broker.
What changed in a release, and roughly what came before it, is in
CHANGELOG.md — the most recent cycle itemized, earlier
history summarized, and the security posture statements worth reading before you
install a mod. Breaking changes stay in wiki/Upgrading.md; the changelog links
to it rather than duplicating it.
Adding or editing launch profiles (WSL / zsh / PowerShell / Git-Bash) is covered
in wiki/Launch-Profiles.md — the recipe catalog, the three
profile fields, the webterm_profiles.json sidecar-vs-broker_config rule, and
the browser-realm-only editing model.
Writing a mod — or installing somebody else's — is
wiki/Writing-a-Mod.md: the trust posture (a mod is not
sandboxed), every mod.json field, the whole ctx surface, CSS ownership, the
help.md + corpus-regeneration rule, and the x- namespace and portable-mod
contract an installable mod has to meet.
The full engineering reference lives in wiki/Technical-Reference.md:
- the complete wire protocol and frame semantics,
- the full auth model (every surface, token precedence, CORS),
- every HTTP endpoint including the MCP HTTP contract and its error table,
- access modes, the MCP + profiles config sidecars, and the shipped MCP server, and
- deployment notes (systemd units, multi-host over Tailscale, testing).
Released under the MIT License — see LICENSE.
not the.primeagen@gmail.com but hi :)