Skip to content

Repository files navigation

image

Browserland

A remote control plane for every machine you own: install Browserland on each one, and a single browser tab becomes a desktop of live terminals across all of them — including the AI coding agents running inside.

Point a browser at one broker and you get a full windowed desktop of live terminals: tile them, tab them, split them, and drag them across virtual desktops. Each terminal is a real PTY running on some machine, streamed to the browser over a WebSocket. Add your other machines as hosts (token-auth'd, e.g. over Tailscale) and their terminals appear in the same desktop — launch a shell on the laptop, the desktop, and the server from one + menu, side by side.

The shells keep running even when no browser is attached. Close the tab, come back tomorrow, and the screen heals from a snapshot — exactly where you left it.

The name says it plainly: a whole little desktop — windows, terminals, and your fleet of AI coding agents — living entirely in a browser tab. (webterm is the Python package and module name.)

Browserland also exposes an MCP server, letting LLM harnesses drive the terminals directly — including full-screen TUIs.

⚠️ Breaking change: a token is now required on every connection

Browserland used to treat loopback as an auth exemption — with no auth_token set (the shipped default), anything on the machine could read every terminal's title/pid/cwd, attach to any PTY, and reach host-wide file read/write. That was never sound: tailscale serve in front of a 127.0.0.1 bind makes every tailnet request arrive from loopback, and any web page you have open can reach loopback too. A token is now required on every route and every interface, with no exemption and no opt-out.

New install? Nothing to do. The broker mints its own token on first run and prints a ready-to-open URL. You never have to create one.

Were you running without a token? Then on your next start:

  • The browser will ask for a password. Get it with python -m webterm.broker --print-token — it's also minted into webterm_token.json beside your state store.
  • Terminals launched by the previous run cannot reconnect and must be relaunched once. Their shells keep running, but they were started without the token and it can't be injected into a live process. This is the one unavoidable loss — finish or accept whatever is in them first.
  • Scripts calling /sessions, /launch or /file/* over loopback now get 401. Send Authorization: Bearer <token>.

Running a fleet? Set auth_token yourself in broker_config.json before the cutover, so every host shares one password you already know instead of a different minted value per machine.

Full details, a does-this-affect-me table, and recovery steps: wiki/Upgrading.md.

What is this?

Browserland is two small programs and a browser:

  • Agents (producers) are headless processes that own a real terminal — pty.openpty on Linux, ConPTY/WinPTY on Windows. An agent runs a command (bash, cmd.exe, an AI coding agent, anything), keeps a ring buffer of recent output, and streams the terminal over a single WebSocket.
  • The broker is a small web server. Agents register with it; browsers connect to it. It relays bytes both ways, serves the desktop UI, and can spawn new agents on demand from a list of pre-approved profiles.
  • The browser renders the desktop — a tiling window manager over xterm.js — and sends your keystrokes back to the PTY.

The wire format is a deliberately small set of JSON frames over one WebSocket, so any producer that speaks them can register with the broker — Browserland's own agent is just the reference implementation.

┌─────────┐  binary ANSI + JSON   ┌────────┐  /ws?session=<id>  ┌─────────┐
│  agent  │ ──────────────────▶  │ broker │ ──────────────────▶ │ browser │
│ (PTY +  │ ◀──────────────────  │        │ ◀────────────────── │ xterm.js│
│ ConPTY/ │  input/resize/       └────────┘  input/paste/resize └─────────┘
│ openpty)│  snapshot_please
└─────────┘

Because the PTY lives in the agent, not the browser, terminals survive browser reloads and even broker restarts — the agent reconnects with backoff and the browser's attach triggers a snapshot redraw from the ring buffer.

One tab, every machine

The part people miss: Browserland is a remote control plane, not just a terminal for the machine it runs on. Run a broker on each of your machines, then register them as hosts in the UI (Control Panel → Hosts — a name, a URL, and that broker's token; a Tailscale address works great). One browser tab then fronts the whole fleet:

                       one browser tab
                             │
          ┌──────────────────┼──────────────────┐
          ▼                  ▼                  ▼
     ┌─────────┐        ┌─────────┐        ┌─────────┐
     │ broker  │        │ broker  │        │ broker  │   one per machine
     │ laptop  │        │ desktop │        │ server  │   (LAN / Tailscale)
     └─┬─────┬─┘        └─┬─────┬─┘        └─┬─────┬─┘
       ▼     ▼            ▼     ▼            ▼     ▼
     shell  claude      shell  build       shell  claude   your PTYs + agents
  • The right-click + menu groups launch profiles per host — open a PowerShell on the Windows box and a zsh on the Linux server from the same menu, and tile them next to each other.
  • Every window is badged with its host; the taskbar's broker status — a chip for one host, a single aggregate badge for several — shows at a glance which brokers are reachable, with per-broker detail in the + menu.
  • Each host keeps its own token, profiles, and file API — the UI fans out, the brokers stay independent. A machine going down takes its windows stale, not the desktop.
  • The MCP server speaks the same multi-host language: one --hosts list lets an AI harness enumerate and drive terminals across all brokers through a single tool surface.

wiki/Setup-and-Onboarding.md walks through joining machines step by step.

Screenshots

The desktop with tiled and floating terminals

image image

Features

  • In-browser tiling window manager (niri-style): floating and tiled windows, a virtual-desktop pager, and a taskbar — all in the page, no native app.
  • Nested tabs and splits: any tile can hold a tab group, and tabs can hold splits, via a recursive cell model. Build the layout you want by dragging.
  • More than terminals: sticky notes, a CodeMirror 6 text editor, a file manager, a task manager, a synced scratchpad, a clipboard history, and a session recorder — each a bundled mod (see Mods) over the broker's token-gated /file/* and per-mod store APIs.
  • Session recording & replay: hit ⏺ on any terminal to record it byte-faithfully, then play it back at the original size — pause, 0.25×–8× speed, continuous rewind, and notes pinned to timestamps.
  • A bundled mod system: the desktop's app windows and widgets are mods over a small, versioned ctx API — a mod is one folder with a manifest and an entry script, and it can add window kinds, per-terminal title-bar widgets, taskbar chips, Control-Panel settings, and in-app help pages.
  • Cross-platform PTY: Linux pty.openpty; Windows auto-selects ConPTY or WinPTY — even headless/detached agents acquire a hidden console and re-enable Ctrl-C so they run ConPTY with a working interrupt and live resize, falling back to WinPTY only if console acquisition fails (or when forced via --pty-backend winpty).
  • AI agent fleet: detects the foreground coding agent in each window (claude / codex / grok / opencode / hermes) and tracks live OSC title + working directory, with an opt-in per-window git-status widget (the git mod).
  • Multi-host — the control plane: attach the same UI to the brokers on all your other machines (e.g. over Tailscale) and run their terminals side by side in one tab, with live broker status in the taskbar and launch menu. See One tab, every machine.
  • Single active-browser lease: exactly one browser drives input at a time, so two open tabs never fight over the keyboard.
  • Opt-in MCP / AI agent access: an MCP client or AI harness can list, observe, drive, and launch terminals — including live interactive TUIs and the console you're working in — under per-window access modes you control. See MCP & AI agent access below.
  • Token auth, no open RCE: one token gates every connection — every route, every interface, loopback included — and doubles as the UI password. Nothing to set up: a broker with no token configured mints one and prints the URL to open. Launching is profiles-only (the client can never supply a raw command).
  • Launch profiles you can edit in the UI: add WSL/zsh/PowerShell shells from Control Panel → Launch profiles (one-click WSL-distro / shell Detect…), applied live with no restart — still profiles-only. See wiki/Launch-Profiles.md for recipes and the security model.

Mods

The desktop's app windows and most of its optional chrome ship as mods — the terminal pipeline, window manager, the multi-host connection model, and MCP surfaces stay core (only the optional sharing of that host list across browsers is a mod). A mod is a self-contained folder under webterm/broker/mods/<id>/ holding a manifest (mod.json), its script(s), and optional CSS + an in-app help page. The bundled ones are trusted first-party code, spliced from an explicit allow-list into the single served page; a broker can also be handed a mod at runtime (POST /mods/install, ids prefixed x-), which is served from its own same-origin URL and picked up on the next page load. Either way a mod registers through a versioned ctx API that exposes per-terminal-window hooks, app-window kinds (which appear in the + launch menu), taskbar chips, Control-Panel settings, a durable per-mod server store with revision history, an in-desktop copy/paste observer, and help cards.

A mod is not sandboxed — it is same-origin code with the page's full authority, and ctx is a set of reviewed choke points rather than a boundary. Writing one (and the trust model behind installing one) is covered in wiki/Writing-a-Mod.md.

A mod's settings sync across your browsers via the broker's shared state; its enable/disable toggle is deliberately per-browser — flip it in Control Panel → Mods. A broker-side mods_enabled master switch gates the whole system, and a broker can pin individual mods on or off for every browser that loads its page. The desktop ships with nineteen:

Mod What it adds Default
theme color-scheme picker for the desktop chrome on
pattern desktop background patterns on
clock taskbar clock with timezone picker on
help the in-app help window + ? chip on
sticky sticky notes on
editor CodeMirror 6 text editor on
file-manager dual-pane file manager on
task-manager live per-host process list on
scratchpad server-backed notes, synced across browsers with revision history on
recorder terminal session recorder + fixed-size player (speed, continuous rewind, timestamped notes) on
host-registry optional shared broker list — publish/pull your host list across browsers on
mousemode title-bar chip while a full-screen app owns the mouse on
mod-sync copy this broker's mod setup to your other brokers (all or selected) on
workspaces vertical workspaces — taskbar pager, per-workspace names, Send to workspace on
git per-terminal git branch + dirty-state widget off
aistatus AI-provider status chip + window off
clipboard rolling history of copies/pastes made through the desktop off
termfont terminal font picker off

clipboard and aistatus are off by default on principle — clipboards carry secrets, and status polling talks to third-party endpoints; git and termfont are simply opt-in preferences. Enabling any of them is one click in the Mods pane.

MCP & AI agent access

The broker exposes a token-gated /mcp/* HTTP API and ships a stdio MCP server (webterm.mcptool), so any MCP client or AI harness can list, observe, drive, and launch terminals. The agents are just producers; the broker stays the sole authority — every MCP call is gated by the same per-window access modes and the master enable switch.

  • Interactive TUIs, as plain text — read_screen renders the current screen of a live terminal by replaying its PTY ring buffer through pyte, so a harness can read full-screen apps (btop, htop, vim, less) — not just line-oriented scrollback — and send_input types into them. Without pyte the read falls back to a dependency-free in-house renderer that still returns a bounded rendered grid (degraded: true is now reserved for a rare last-ditch raw decode).
  • The live session you're working in — list_terminals enumerates running sessions (id, title, cwd, agent, kind, cols/rows, mode), so a harness can attach to the exact console a person is using right now, read its state, and (in readwrite) drive it. Sessions persist across browser reloads and broker restarts, so the handle stays valid.

Tools

Each tool maps to a broker endpoint and returns its JSON. Window ids are namespaced "<host>:<int>" strings so one server can front several brokers (see Multi-host below); with a single broker the host is default ("default:12345").

Tool Endpoint Notes
mcp_info(host?) GET /mcp/info feature flags (allow_launch, default_mode). Omit host → dict keyed by host name
list_terminals GET /mcp/terminals {"terminals":[…], "errors":{host:msg}}: all hosts merged, each terminal's host set to the config name (broker's machine hostname preserved as machine_host) + namespaced id; a down host lands in errors without sinking the rest. Each terminal carries a pyte flag — false = the agent lacks pyte, so attr_runs/keyframe-repair are unavailable and sparse alt-screen frames are flagged partial only (#134)
list_profiles(host?) GET /mcp/profiles launchable profile names + default. Omit host → dict keyed by host name
read_screen(id) POST /mcp/read screen rendered as a bounded plain-text grid (pyte, or a dependency-free fallback); result carries content_hash, stable_hash (the cursor-blind digest — a cursor blink in place doesn't change it, a cursor move does), and idle_ms (best-effort ms since the last PTY output — absent from older agents, unreliable for a perpetually-animating app). Waits (one call, timeout_ms-bounded, mutually exclusive): wait_for_change / wait_for_text / wait_for_regex / wait_for_idle (block until the screen settles — stable_hash unchanged for N ms)
send_input(id, data) POST /mcp/input target window must be in readwrite mode; newlines are sent as Enter (CR) so commands run (incl. on PowerShell)
send_keys(id, keys, delay_ms?) POST /mcp/input send control/escape keys — ["C-c"], ["Esc"], ["Up","Enter"] — that plain text can't express; delay_ms (or a per-terminal set_pace default) writes one key per POST for a frame-polling TUI
set_pace(id, pace_ms) POST /mcp/pace readwrite; set a per-terminal DEFAULT send_keys pacing (ms, capped 1000, 0 disables) so multi-key sends auto-pace without passing delay_ms — for a frame-polling TUI (Dwarf Fortress). Broker-local + ephemeral (resets on agent reconnect)
reset_terminal(id) POST /mcp/reset readwrite; correlated round-trip that wipes the agent's screen-render buffer so the next read_screen starts clean (502 on a non-agent producer)
flush_input(id) POST /mcp/flush readwrite; correlated round-trip that discards keystrokes queued to the app but not yet consumed — the input-side mirror of reset (502 on a non-agent producer; a no-op on a Windows/ConPTY agent)
launch_terminal(profile?, cols=80, rows=24, title?, cwd?, host?) POST /mcp/launch broker must have allow_launch enabled; host required when multiple hosts are configured

Multi-host. Pass --hosts (or $BROWSERLAND_MCP_HOSTS) a JSON array of {name, url, token} descriptors to serve N brokers from one server process; every id-taking tool routes on the "<host>:…" prefix. The single --broker-url/--token form is the one-host shorthand (default). See webterm/mcptool/README.md for details.

Scopes. Several MCP clients can share one broker without seeing each other's windows: a server started with --scope projA (or BROWSERLAND_MCP_SCOPE, or a per-host "scope") sees only the windows tagged projA, while a server with no scope keeps the admin view of everything. A window that server launches is tagged for it and starts in readwrite; tag a hand-started one from its robot button or title-bar menu. Control Panel → Access → MCP access → Copy .mcp.json writes a ready-to-save client file for a scope. A scope is a convention, not security (every client holds the same token), and a window's scope and MCP mode now survive a broker restart. See wiki/MCP-and-AI-Agents.md.

The send_input tool maps newlines in data to a carriage return — the byte a real Enter key sends — so a line submits on PowerShell/PSReadLine (which treats a bare line-feed as a soft continuation) and on a Unix shell alike. The raw POST /mcp/input endpoint stays verbatim: drive it directly to send a literal LF or hand-crafted control/escape bytes.

Safety / enabling

Access is layered and opt-in — nothing is reachable until you turn it on:

  • Master enable is off by default; while off, every /mcp/* call returns 403 mcp_disabled.
  • Per-window mode is off / read / readwrite, with a global default_mode for new windows (a scoped launch_terminal starts in readwrite unless it passes a mode, and a mode set on a window survives broker restarts). off hides a window entirely; read allows observation; readwrite additionally allows send_input.
  • allow_launch is a separate gate for launch_terminal.
  • The MCP token is a bearer secret distinct from the browser auth_token (the UI password): the broker pins it with WEB_TERMINAL_MCP_TOKEN (or the webterm_mcp.json sidecar), and the MCP client passes the same secret via BROWSERLAND_MCP_TOKEN (as the harness examples below show).

Register with a harness

claude mcp add browserland \
  --env BROWSERLAND_MCP_TOKEN=… \
  --env BROWSERLAND_MCP_URL=http://127.0.0.1:4445 \
  -- python -m webterm.mcptool

Any other stdio MCP client (Hermes, your own, …) registers the same way — point it at the launch command and pass the two env vars:

{
  "mcpServers": {
    "browserland": {
      "command": "python",
      "args": ["-m", "webterm.mcptool"],
      "env": {
        "BROWSERLAND_MCP_TOKEN": "…",
        "BROWSERLAND_MCP_URL": "http://127.0.0.1:4445"
      }
    }
  }
}

Or run the server directly, talking to the local broker:

BROWSERLAND_MCP_TOKEN=… python -m webterm.mcptool
BROWSERLAND_MCP_TOKEN=… browserland-mcp --broker-url http://127.0.0.1:4445

For the full HTTP contract, error table, and config sidecar, see wiki/Technical-Reference.md and webterm/mcptool/README.md.

Quick start

You need Python ≥ 3.9. Install from a checkout:

pip install -e .

Windows

# broker (default 127.0.0.1:4445) — mints a token on first run and prints
# the URL to open, token included
python -m webterm.broker

# the token again, any time (never mints one):
python -m webterm.broker --print-token

# an agent running cmd.exe, registered with the local broker
python -m webterm.agent -- cmd.exe

# then open the printed http://127.0.0.1:4445/?token=... and click the
# session (or "new terminal")

Windows agents also need a PTY backend: pip install -e ".[windows]" (pulls in pywinpty).

Linux

./launchers/run-broker.sh                  # broker
./launchers/run-agent.sh -- bash -l        # agent

The launcher scripts bootstrap a virtualenv with the runtime dependencies on first run. The broker prints the URL to open, token included; python -m webterm.broker --print-token reprints it later. run-agent.sh picks the token up from webterm_token.json automatically, so a hand-started agent on the same box needs no setup.

Install & extras

pip install -e .                       # core: broker + agent
pip install -e ".[windows]"            # + pywinpty (Windows PTY backend)
pip install -e ".[pyte]"               # + pyte (tier-2 snapshot rendering)
pip install -e ".[procs]"              # + psutil (task manager, agent badge, live cwd)
pip install -e ".[mcp]"                # + the stdio MCP server (Python ≥ 3.10)
pip install -e ".[dev]"                # + pytest for the test suite

psutil (the procs extra) is best-effort: it powers the task-manager process list, the foreground-agent badge, and live-cwd tracking. Without it the agent still runs and still destroys windows — those three views just degrade (empty list / no badge / no cwd).

Mix and match, e.g. pip install -e ".[pyte,mcp,dev]". The mcp extra requires Python ≥ 3.10 (the MCP SDK), while everything else runs on Python ≥ 3.9. See MCP & AI agent access for running the server.

Project layout

Path What
webterm/agent/ Headless producer: PTY backends, output ring buffer, OSC-title sniffer, reconnecting WebSocket client
webterm/broker/ Web server: desktop UI (ui.py assembles the served page from ordered *.html/*.css/*.js fragments), /ws relay, producer WS, session list, profiles-only launch
webterm/broker/mods/ The bundled desktop mods — one folder per mod: mod.json manifest, script(s), optional CSS + in-app help page (authoring guide: wiki/Writing-a-Mod.md)
webterm/mcptool/ The shipped stdio MCP server wrapping the broker's /mcp/* API
webterm/protocol.py The single source of truth for the JSON frame shapes
launchers/ venv-bootstrapping run scripts (and systemd units) for both OSes
tests/ pytest suite: protocol, snapshots, agent↔broker integration, real-PTY round trips

Documentation

Everything lives in wiki/ — one tree, and the same source the in-app Help window reads. Open Help with the ? chip in the taskbar to search it without leaving the desktop; the developer and operator pages are in there too, behind Include developer docs.

Page
Start here Home · Getting Started · Keyboard Shortcuts
Building layouts Window Modes · Arranging Windows · Columns & Widths · Snapping & Pop-out · Floating Window Controls
The desktop Workspaces · Taskbar · Context Menus · Window Types
Multi-host & AI Hosts & Multi-Browser · MCP & AI Agents
For developers Setup & Onboarding · Launch Profiles · Writing a Mod · Technical Reference · Upgrading

New here, or setting up more than one machine? Start with wiki/Setup-and-Onboarding.md — the onboarding guide: the broker / agent / browser mental model, joining machines over Tailscale via Control Panel → Hosts, what not to hand-edit, and running the broker unattended in the background (Windows Task Scheduler / Linux systemd). It's written to be followed by a human or a coding agent.

Upgrading an existing install? wiki/Upgrading.md is the breaking-change record — what changed, whether it affects you, and how to recover. Most recently: a token is now required on every connection, including loopback (#142), which strands terminals launched by a previously tokenless broker.

What changed in a release, and roughly what came before it, is in CHANGELOG.md — the most recent cycle itemized, earlier history summarized, and the security posture statements worth reading before you install a mod. Breaking changes stay in wiki/Upgrading.md; the changelog links to it rather than duplicating it.

Adding or editing launch profiles (WSL / zsh / PowerShell / Git-Bash) is covered in wiki/Launch-Profiles.md — the recipe catalog, the three profile fields, the webterm_profiles.json sidecar-vs-broker_config rule, and the browser-realm-only editing model.

Writing a mod — or installing somebody else's — is wiki/Writing-a-Mod.md: the trust posture (a mod is not sandboxed), every mod.json field, the whole ctx surface, CSS ownership, the help.md + corpus-regeneration rule, and the x- namespace and portable-mod contract an installable mod has to meet.

The full engineering reference lives in wiki/Technical-Reference.md:

  • the complete wire protocol and frame semantics,
  • the full auth model (every surface, token precedence, CORS),
  • every HTTP endpoint including the MCP HTTP contract and its error table,
  • access modes, the MCP + profiles config sidecars, and the shipped MCP server, and
  • deployment notes (systemd units, multi-host over Tailscale, testing).

License

Released under the MIT License — see LICENSE.

not the.primeagen@gmail.com but hi :)

About

Browserland is a tiled multi server terminal in browser with MCP

Resources

Contributing

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages