Cluster, step 1 of 5: a gateway in front of independent daemons - #103
Merged
Merged
Conversation
…t until the next garbage collection acquireSingleWriterLock's handle was discarded at boot. better-sqlite3 closes a handle whose object is garbage collected, and a closed handle releases the exclusive file lock — so the "one ledger, one daemon" guarantee lasted only until the first collection: measured 2026-09-11 with the fake executor, a second daemon on the same ledger started and listened seconds after the first. The handle is now kept in main.ts and closed on shutdown; the second instance dies at boot naming the conflict, as intended.
…paths the gateway imports without the executor The gateway in front of several daemons wants four small self-contained pieces of the daemon — the constant-time token compare, the per-name serialization queue, the single-writer file lock and the bounded shutdown — and nothing else. Four subpath entries expose exactly those; deliberately not the package root, whose import graph loads dockerode, execa and the AWS SDK into any process that touches it. The lock's busy sentence becomes the caller's: the gateway takes the same lock over its own database file and must name its own variable in the exit. auth.ts reads the cookie jar structurally so it type-bundles on its own, without @fastify/cookie's request augmentation in its graph.
…id, inside the name's slot when a create is in flight The gateway keeps no directory of where sandboxes live — a second copy of a fact the nodes' ledgers already hold, and every copy drifts. When its cache has no answer it asks every node this one read-only question in parallel and routes to the node that says yes. Three steps make the answer truthful about a create in flight. A row that exists answers at once, whatever its state: a restoring sandbox has a row, and waiting for its slot would hold the answer for the whole restore, long past the gateway's two-second patience. No row while the name's slot is busy waits its turn and looks again — the daemon creates first and writes the row second, both under the slot, so a gateway retrying a create whose answer was lost finds the sandbox on the node that built it and never places a second copy. No row and a free slot is a plain no. By id there is no slot to wait on and none is needed: nobody can ask about an id before the create that minted it has answered. The in-flight test parks the fake executor's create to hold the acquire mid-build; removing the wait turns it red.
…nd where it can be reached DORMICE_GATEWAY_ENDPOINT makes a daemon a node of a fleet. Set, it POSTs /checkIn to the gateway every DORMICE_CHECK_IN_INTERVAL_SECONDS with its id, the address the gateway may forward to (DORMICE_NODE_ENDPOINT, defaulting to its own loopback — right when gateway and node share a machine), its build and a fresh reading: CPU, memory, swap, the data disk, the ledger's census by state. Unset, the daemon is the whole platform by itself, as before. The token is the one DORMICE_API_TOKEN gateway and nodes share. Who knows the truth speaks: the node knows what it runs and where it lives, so the node reports and the gateway only listens. The gateway learns of a node from its first check-in — no registration verb, no nodes file — and reads two missed check-ins as down; the interval travels in every check-in so both ends measure with the same number. Failures are logged on the change, never every tick, and are never fatal: the gateway is the fleet's front door, not the node's reason to live. The shared host schema is split into its named parts (host reading, data disk, state counts) so getHostMetrics and the check-in describe the machine with one vocabulary; readHostReading is the one function both read through.
…e placed, existing ones are found by asking, everything else is forwarded raw A fleet is N daemons that know nothing of each other behind one gateway. The gateway holds no sandbox state: it keeps the nodes that have checked in (one table, written by the nodes themselves — no registration verb, no nodes file), what each last reported (memory), a cache of where names were last found, and a per-name slot. When the cache has no answer it asks every node "do you hold this sandbox?" in parallel, two seconds, and routes to the one that says yes; two yeses are a 409 it refuses to guess about; no yes while a node is silent is a 503 with Retry-After, because a name that lives only on a silent node must not be built a second time elsewhere. A new name is placed on the emptiest node by active density per core, skipping nodes that are down, above the CPU limit, at the active ceiling or below the data-disk floor, with each pick counted against its node until the next reading. The forwarding plane, error dialects and placement come from the first cut by file (branch router-v1), minus the directory, the claims, the reconciler and the actor header: one token for the whole fleet, and the node trusts the gateway whole. Faces in this cut: the native sandbox verbs, the E2B control plane by name or id, envd by header, the bare signed-URL door as an honest 501, and the gateway's own checkIn / listNodes / removeNode. Daemon-addressed verbs answer 501 until the configuration authority moves here; the sandbox port proxy joins with the sandbox domain. Only the fleet token opens the door in this cut; minted keys arrive with the key table. Reverse-proved: without the name slot twenty simultaneous acquires build several copies; without the placement counter a burst lands on one node; without the 404 re-check a destroy behind the gateway's back leaves a stale entry.
… holds a placement slot The placement counter exists so a burst inside one interval is counted against a node before its next reading shows it. It was only ever reset by the reading, so a sandbox placed and destroyed within the same interval — a short job, or the exam's churn — kept holding a slot the reading would never show; with the exam's two-per-node gate that shut placement for a whole interval after a handful of fast tests. Each node now remembers the ids its creates answered with since the last reading; a destroy of one of them takes the placement off the count, and a create the node itself refused is uncounted at once. A hop's 502/503/504, or no answer, still counts until the reading: the sandbox may exist.
…at joins and dies, all driven over the wire The setup boots node A exactly as before (every existing suite is untouched) and, outside docker mode, a gateway plus nodes B and C that share its token and check in every second. The exam speaks only the SDK, the official e2b package and plain fetch; direct calls to a node stage what the gateway must then find: a sandbox built behind its back is routed at once (no reconcile to wait for), a name on two nodes is a 409 naming both, a destroy behind its back is caught by the 404 re-check. Placement, the active gate and its Retry-After, five simultaneous acquires, streaming through two hops, envd and the bare signed door are graded; a node booted by the test itself joins at its first check-in, and when it dies its sandboxes 502 and a new name is a 503 naming it until removeNode. A second gateway on the same database file dies at boot naming the conflict, as does a second daemon on node A's ledger.
…ngeset for the fleet wire The unit is copied into place by hand until install.sh learns the two roles; it carries no docker dependency, because the gateway never touches a container.
… at the door, and a check-in that moves or shares an endpoint is said in the log forwardCreate took every 2xx id into placedIds and paid every node-side 4xx back to placedSinceCheckIn, wakes included. A wake was never counted in the first place, so destroying one inside the same interval uncounted a real placement and let the active gate admit one sandbox more than the node's reading allows. The callers now say whether pick() chose the node (a placement) or the name was found on it (a wake); only a placement moves the count either way. Reverse-proved: without the flag the new test reads b's count as 0 with a placement still in flight. nameOf and the E2B metadata.name are judged by the shared sandboxNameSchema before any node is asked. A node's 400 to lookupSandbox read as silence, so a 129-character name came back as a 503 with Retry-After — retry forever — instead of the 400 it deserves. Fleet.checkIn reports the endpoint a node moved from; the check-in route warns on a move and when two nodes report one endpoint. Both are misconfigurations (two machines sharing DORMICE_NODE_ID, whose default is node-1; a DORMICE_NODE_ENDPOINT naming the wrong machine) whose symptom downstream is a 409 on every name, and the check-in is the only place that sees them. e2b.ts: the `dialect` helper that existed only to be overridden is gone.
…OINT: the loopback default names the gateway's own machine Left unset, the node reports http://127.0.0.1:<DORMICE_PORT> as where it can be reached, and a remote gateway dials its own daemon (or nothing) for this node: every sandbox placed "here" lands there, and every lookup then finds it twice — a 409 on every name, from one forgotten variable. Refused at boot, naming the variable and the address to write, whenever DORMICE_GATEWAY_ENDPOINT is not loopback. An explicit value is the operator's word and is taken as written. The loopback test tolerates an unparsable URL: zod runs the object-level refinements even when a field failed, and the field's own error ("must be a full http(s) URL") must be the one that shows, not an Invalid URL thrown from inside the rule.
…swers at once when the slot is free tryRun-then-run did the same thing in three lines. run() executes the task immediately when nothing holds the key and queues it behind the holder otherwise — exactly "a plain no when the slot is free, wait for the acquire when it is not". SKIPPED is no longer imported here.
…nstead of followed The gateway hands a node's endpoint to undici as the request's origin when it forwards, and undici refuses an origin that carries a path (UND_ERR_INVALID_ARG, measured). Its own lookup joins `<endpoint>/lookupSandbox` as a string and works with a path. So a node reporting http://10.0.0.7:80/dormice was found by every lookup and reached by no forward. The shared wire now has `endpointSchema`: a trailing slash is dropped (the two ends must agree byte for byte, and `//lookupSandbox` is a 404 the gateway reads as silence), a path or query is a 400 at /checkIn, and DORMICE_NODE_ENDPOINT refuses the same at the daemon's boot, where the operator is looking. The node's check-in used fetch's default redirect handling. A front that answers plain http with a 308 to https (a Caddy binding the gateway's domain) would be followed across origins, and the Fetch standard drops Authorization on the way, so the gateway answered 401 and the log said "wrong token" where the address was wrong. The check-in no longer follows: it reports the 3xx and its Location, and names the variable to fix. Found by the second review pass of 2026-09-14; each fix reverse-proved (the new test goes red with the fix removed).
forwardCapture opened with `if (res.destroyed) return null`; forwardStream did not. A caller of a streamed verb (execCommand) that hangs up during the lookup round — up to two seconds when a node is slow — has a destroyed response whose 'close' has already fired, so the abort wired to 'close' would never come, and the node would run the command to its end for nobody. Same first line now, and the doc comment that already claimed the two followed one rule is true. Found by the second review pass of 2026-09-14; reverse-proved.
…d a second reporter under one id are refused while the node is live Three ways the fleet could end up with one name on two nodes, or a placement past every gate, found by the second review pass of 2026-09-14. The cache is a hint for every verb but the two creators: a stale entry costs a reader one 404 that evicts it, but acquireSandbox and the E2B named create are create-or-wake on the node, so a cache hit forwarded as a wake to a node that has since deleted the row (an E2B deadline kill is the scanner's routine, five minutes by default) rebuilt the sandbox there — past placement's four gates, uncounted in placedSinceCheckIn, every re-create of an expired name pinned to its first node. The creators now confirm a cache hit with that one node by id (Finder.byName confirm): found is a wake, absent evicts and asks the fleet, silent is the same 503 a silent stranger earns. One RTT per warm acquire. removeNode on a node still checking in deleted its row; a name of its acquired before its next check-in was placed elsewhere, then the node re-added itself and the name was on two nodes — a 409 an operator clears by hand. It is now a 409 up front, naming what to do: stop the daemon, wait two of its intervals (what "down" means), then remove. A check-in that changed a node's endpoint inside the previous reporter's own interval was written through. Two daemons sharing one DORMICE_NODE_ID (the default is node-1) flipped the endpoint at every check-in; a lookup asked whichever was current, a name on the other read as new and was built again, on two nodes, with no 409 ever. Such a check-in is now refused (409) naming both addresses; the first reporter keeps the id; a node that really moved is taken an interval later. The 503 for a silent node now says the node may be busy building that very name (its lookup waits for the name slot, and a cold create outlasts the gateway's two seconds), and the shared lookup comment no longer promises that a retry finds a sandbox still being built.
Node holds a written head until the first body byte. A node stream that opens and then waits for its first event looked, through the gateway, like a node that had not answered — measured: a 300ms pause before the first byte held the status line for 300ms. The head now goes out when the node's did.
…over a few milliseconds The sampler was primed right before the first check-in, so its first delta spanned the sliver between the two calls and read near 0 or near 100 by luck; a freshly restarted node could sit out its first interval on a number that meant nothing. Unprimed, the first reading is null — "no interval yet" — which placement lets through as unknown, as it was written to.
…s why in words undici's request() ignores an abort signal while the socket is still connecting, so a node whose host drops the SYN (a deleted VM, a closed security group — the case removeNode exists for) held every question for the dispatcher's 10s connect timeout, not the promised two seconds; and the creator's confirmation runs inside the name's slot, so queued acquires of a name cached there paid that in series. fetch honours the signal mid-connect (measured: request 10 500ms, fetch 2 001ms against 192.0.2.1); it follows no redirect here either. causeOf now prefers string codes: a DOMException's numeric legacy code rendered a timed-out node as "did not answer (23)". One causeOf, in lookup.ts, shared with forward.ts. The refused-connection test moved off port 9, which fetch refuses as a "bad port" without dialling.
clientGone was checked only at placement, after the lookup round: a creator queued behind a slow create or destroy whose client had given up still confirmed with the cached node — a question inside the slot, paid in series by every abandoned request behind a stuck name. It is now the first thing inside the slot on both faces, and still checked after the round of questions. The fake node learned a slow destroy for the test.
…talled gateway unit first A machine that runs both (the single-machine install is a fleet of one) upgraded with install.sh rebuilt only the server, CLI and console: the gateway's dist stayed on the older commit and ran it at its next restart, against a daemon whose check-in it may no longer parse. The gateway is now in the build filter, and a dormice-gateway unit that is running or enabled is restarted before the daemon so its first check-in lands on the new one.
The non-200 branch cut the gateway's answer at 200 characters. The gateway's longest refusal, the 409 naming both endpoints of a shared node id, runs to about 260 — the operator saw the diagnosis and lost the remedy. Cut at 400 now: whole enough for every sentence the gateway writes, short enough that a front's HTML error page does not flood the log.
It claimed the daemon runs its other verbs unserialized; the daemon holds a slot per name for thirteen of the fifteen. The gateway's own reason for taking the slot only for acquire and destroy — the two verbs whose outcome it acts on — stands on its own and is now the one stated.
A node deletes rows on its own (an E2B deadline kill is the scanner's routine) and tells no gateway, so the cache kept an entry for every sandbox ever created through the process for as long as it lived — an unnamed E2B create a minute is half a million dead entries a year. A hundred thousand entries, generous next to a fleet's live population; past it the entry nobody asked about for longest goes, and a hit moves an entry to the young end. A wrongly evicted entry costs the one round of questions any miss costs.
…names the way out A node whose DORMICE_NODE_ID changed checks in as a new node while its old id keeps its row; both are asked, both answer from the same address, and every name there was a 409 telling the operator to destroy one copy — of a sandbox that exists once. The finding now carries the endpoints, and when they are one the sentence says so: remove the id that no longer checks in (or correct a DORMICE_NODE_ENDPOINT that names the wrong machine). Nothing is collapsed or healed; the fleet's list is wrong and the operator is told exactly how.
…hen the failure changes fetch says "fetch failed" and keeps ECONNREFUSED, ENOTFOUND or the TLS error in cause; the log carried only the first half. And a failure was logged once per streak: a gateway that was unreachable and then, up again, refused this node as a twin of another (409) never made the log — the streak had already been announced. What is wrong is now compared with the numbers blanked, so a change of failure is one more line and a 409 that says 3s ago, then 4s ago, is still one.
…_NODE_ID The gateway tells nodes apart by it, and node-1 is what every other unconfigured node says too: the second to check in was refused as a twin at every check-in, or — when the first had been silent for an interval — taken for it having moved, and the first's names placed again elsewhere. The same rule as DORMICE_NODE_ENDPOINT, for the same reason: refused at boot, where the operator is looking. Beside its gateway, or alone, the default still serves.
… for thirty seconds, and a shared endpoint is warned about once Two leftovers from the third review. After a gateway restart every node is silent so far, the running ones included, and downReason alone let an operator remove a live node in the first interval; the Fleet now records when it started and removeNode refuses a never-heard node for two default check-in intervals, naming the wait. Two nodes reporting one endpoint were warned about at every check-in, two hundred and forty lines an hour; the route now remembers what it last said per node and speaks when the situation arises, changes or ends.
…e each, not rows in a table nobody queried Design record #16 (2026-09-13). The bounded activity table, the listActivity verb, the SDK method, the shared ACTIVITY_KINDS and actor vocabulary, the console's activity page and the workbench's activity card (with its message domain in all ten locales) are deleted, and migration 0023 drops the table. What used to be an event is now a line in the daemon's own log where a logger already is: the routes log the sandbox-addressed changes, the key management and the settings, ingress and upgrade writes through the request log; main.ts logs the start and the disk growth; the heartbeat's reconcile and scan summaries were already logged. The lifecycle engine, the reconciler and the archiver stay silent and no longer carry an actor parameter — attribution existed only to feed the ring, so the auth hooks now answer yes or no and the request carries no identity. A shell death is still written to the row as lastExit, which the wire and the console read; the FakeExecutor records shell removals so the tests that used to count 'rebuilt' events still see a swap happen, or not.
Design record #23 (2026-09-13). DORMICE_MAX_SANDBOXES, the maxSandboxes settings column and knob, the two count-based 429 gates at acquire and E2B create, the capacity figure in getHostMetrics, the console's capacity dialog and the overview's total-against-cap card are deleted; migration 0024 drops the column. The number counted every row — active, frozen, stopped, archived — and none of those is a physical ceiling: stopped rows cost only their disk, archived ones nothing local. The ceilings that exist each have their own reading: the data disk's free space, the host's CPU and memory, which the gateway places by and the operator watches. Both production machines had already set the cap to 100000 on 2026-07-19 to take it out of the way. The overview's third card now shows the total with the three cold-state counts under it; the SDK, e2e and settings tests that used the knob as their "any knob" example move to pidsLimit.
…sion, templates, API keys, the console account The first half of moving the configuration authority off the daemon (design record #22). The gateway's database gains the settings row — seeded once from the env in the daemon's own variable names, with a version that counts every change the nodes must hear about — the templates, the API keys and the console account, and the nodes table gains the one per-node knob, the managed swap target. Migration 0001. Nothing serves them yet: the verbs, the gates and the console arrive in the next commits, the node pulls the bundle after that. The daemon exports its S3 store as a subpath so the gateway's probe can use the real client, and the archive default moves to the shared policy module, where both seeds read it.
…re it restarts, imports the old ledger once, and re-points the right Caddy file The role is not a file: /etc/dormice/env names the gateway, and a remote one makes a node machine. The first install of a node takes --role node --gateway <url> (--node-id, --node-endpoint optional; the token from the environment, never a flag); re-runs on either role take no flags. A node machine gets the daemon alone, a Caddy on :80 to it with no marker and no source-IP gate (the security group is the fence), learns the fleet's registry and base image from the gateway's getConfig, pins the registry's certificate on first sight with its fingerprint printed, and pulls the base image for doctor's probes. The gateway's machine gets the fleet registry: distribution's static binary, pinned and checksummed, as a systemd unit over TLS (a self-signed ten-year certificate, SAN the listening address) with an htpasswd whose one user takes the fleet token — a registry over plain HTTP cannot take basic auth, so the choice was a lock with TLS or no lock. Docker trusts the certificate through certs.d, no restart. The base image is pushed once; its tag and the registry's address become the gateway's late seeds (appended to an existing gateway.env), and a new env file no longer carries DORMICE_BASE_IMAGE. Before any unit restarts, both databases are snapshotted with SQLite's online backup API (three kept). When the gateway's database does not exist and the daemon's ledger holds a settings row, import.js carries the old configuration over before the gateway's first start; a failure ends the run with nothing restarted. The Caddy re-point targets the file the gateway owns (DORMICE_INGRESS_FILE), only under its marker, and any other file under /etc/caddy still aimed at the daemon is named with its line, not touched. Every curl that needs the token reads it from stdin. dor doctor says the base image is the fleet's setting when the env lacks it, names the registry pull as the fix when there is one, and gains a check that the fleet registry answers over the pinned certificate.
…es and their image verdicts; getConfig carries the two new settings
…images through the fleet registry, the two new doctor checks
…s tell fulfilled, the gateway's own upgrade the remedy
rolling.ts judged any build other than the gateway's as behind, so a
node whose build was newer — a commit that landed on main after the
gateway's machine upgraded and before that node's turn came, or
install.sh run on the node by hand — was told to upgrade, pulled the
main head it already ran, and twenty minutes on read stuck with a
remedy (tell it again) that repeated the mistake. Found by review,
reproduced with the pure functions.
Behind now means older: the commit's time orders the two (main is
trunk-based and linear). A newer build is `ahead`: never told,
unreachable when silent like a current one, refused by applyUpgrade
{nodeId} with the reason, and its check-in fulfils any standing tell.
Two commits in one second tie and read behind — the one misjudgement
left, costing a rebuild. The eighth state reaches the wire enum, the
version card (amber, no button), its ten locales and the upgrading doc.
…he gateway machine's node stops before the gateway restarts, a failed push or pull logs out, and a moved registry address is refused import.js creates the gateway database (its migrations) before it reads the ledger, and install.sh guards the import on that file's existence: a failure past the migrations left a file with tables and no settings row, a re-run skipped the import, the gateway seeded from the env, and the daemon's next boot dropped the tables the import carries — keys and console account gone, the backup directory the only copy. Found by review, reproduced with the built tool. The file did not exist before the step: on failure it is removed, and the re-run imports again. The gateway restarted first with the old daemon still running. In the second or two before its own restart, the daemon's check-in could land on the new gateway, read as behind and be told to upgrade — into the very unit that was running (one chance in eight per upgrade at a fifteen-second interval). The daemon now stops first and starts last; a gateway that does not answer /healthz has the daemon started again before the run dies. docker login was followed by a push (gateway) or a pull (node) whose failure exited under set -e before the logout, leaving the fleet token in /root/.docker/config.json: both log out on failure too. --registry-addr on a re-run moved the listener while the env line and the settings row kept the old address, every node pulling from where the registry no longer was: a differing flag is refused, as --gateway is against a node's env. The pull hint (executor, templates doc) names docker login before tag and push, and the upgrading doc says what install.sh does edit in the env files.
…re-tell too; --registry-addr is refused on a node; the tie comment says where it can and cannot bite Rolling.onCheckIn cleared the operator's re-tell only together with a tell on the row. A node re-told while behind, then upgraded by hand to a build ahead of the gateway's, kept its re-tell in memory: the moment it read behind again — the gateway upgraded past it — it was told at once, whatever the one-at-a-time rule said, and the fleet lost two nodes' sandboxes together. The re-tell was for the node that was; a check-in off the old build spends it. --registry-addr on a node machine was read by nobody: the registry is the gateway machine's, and a node pulls from the address its gateway names. Refused before anything is installed, as the other mistaken flags are. The import's cleanup removes the lock's journal file too. The ahead/behind tie: the comment claimed same-second commits are a rebase's rarity. They are not rare in this history (24 of main's last 300 commits share a second with a neighbour), but the two builds compared were each built at a branch head, and two heads a second apart would be two pushes a second apart — a hand-built mid-series checkout is the one way to the misjudgement, and it costs that node one rebuild.
…ndpoint are refused on the gateway's machine and when they contradict a node's env file
The operator's re-tell (applyUpgrade {nodeId}) lived in Rolling's memory
until the node was told or checked in off the old build. A node re-told,
then removed, left it standing: a machine re-imaged under the same id
joined as a new node and was told at its first check-in, past the
one-at-a-time order. Removal forgets the hand along with the row.
--node-id and --node-endpoint were read only at a node machine's first
install. On the gateway's machine, or on a node's re-run with a value
that contradicts /etc/dormice/env, they were taken for a change and did
nothing. Refused before anything is installed, with the edit that does
change it, as --registry-addr and --gateway already are.
…elling it past the order; install.sh judges the four fleet flags against the role in one place
applyUpgrade {nodeId} used to mean "tell this node at its next
check-in, whatever the one-at-a-time order says", kept as a set of node
ids in the gateway's memory until the tell. Three reviews in one day
each found a way for that memory to outlive the node it was for — a
check-in that read current or ahead, the node's removal — and tell two
nodes into one minute. The roll promises one node down at a time; a hand
that could break the promise was the wrong hand. Now the hand forgets a
stuck node's tell on its row: the node reads behind and is told at its
turn, after the node upgrading now, never beside it. Refused everywhere
else in words — 409 on a node upgrading, whose tell is what the rule
counts; 400 on one behind, in line already. Rolling keeps nothing of its
own: every verdict is a function of the rows, the gateway's build and
the clock. The console's toast, the SDK's and the docs' wording follow.
install.sh judged each fleet flag where it was consumed, hundreds of
lines and a role branch apart — and each review found one flag the
other role's machine, or a re-run with a contradicting value, silently
ignored. The four (--gateway, --node-id, --node-endpoint,
--registry-addr) are now judged once, at the top, with the role known
and nothing installed: the other role's flag is refused; a flag that
contradicts the env file's line is refused with the edit that does
change the value; one that repeats it is harmless.
…node it was sent to; Rolling.retell is unstick; the docs stop calling the console's overview dark and the list verbs unrouted The node's applyUpgrade took the gateway's nodeId field and dropped it, then launched its own install.sh — shared upgrade.ts had said the field is refused on a node since the fourth cut, and nothing refused it. Now a 400 that names the gateway's verb; the fake executor's 400 still answers the bare call (app.test pins the order). The hand on a stuck node forgets its tell and never tells again since 9537557; the method was still called retell after the semantics it replaced. unstick, in the gateway and the console. console.mdx said the overview and the sandbox list stay dark until the gateway answers those verbs (it has since the third cut), that a node "has not reported since the gateway started" (the grace went with the fourth cut's rows), and that the workbench shows lifecycle events (the activity table went with the second); quickstart.mdx sent the list verbs to the daemon's port. The SDK's nodeId doc gains the node's answer.
…ht could pulse once more when the file had grown meanwhile
…way's own loopback port; install.sh re-points Caddy to the gateway only once the gateway answers The gateway listens on 127.0.0.1 only, like the daemon, so the `--gateway http://<gateway machine>:3677` the docs gave a node's first install would have died at the fleet-join step. The docs, the installer's own usage text and its closing hint now name :80 — the gateway machine's Caddy — and the join failure says why. The installer used to re-point the gateway machine's Caddy from the daemon (3676) to the gateway (3677) in the ingress step, before the gateway was ever started: on a machine moving to the gateway, the public API face would proxy to a dark port from that reload until the gateway's first start — past the registry install, the base image push, the backups and the import, minutes on a production ledger. The re-point now happens in the services step, after the gateway answers /healthz and while the daemon is stopped anyway; a gateway that does not come up leaves the door on the daemon, which is started again. README: a fleet is a current feature, sharded rather than distributed; "one daemon per machine" is the invariant, not "one machine".
…d under --mirror cn takes better-sqlite3's prebuilt binary from npmmirror better-sqlite3's install fetches a prebuilt binary from GitHub and, failing that, compiles — for which node-gyp fetches the running Node's headers from nodejs.org. On a fresh cn-beijing VM both timed out and the fresh install died in the build (2026-09-16). The headers are inside the Node tarball the script unpacks into /opt, so node-gyp is pointed there whenever that Node is the one running; with --mirror cn the prebuilt binary comes from npmmirror's copy of the GitHub releases, so the compile is not needed at all. A host whose own Node passed the version check has no headers under /opt and fetches as before.
…mirror cn installs the same packages from USTC's mirror when the script's mainland mirror fails
Two things failed on fresh mainland VMs an hour apart (2026-09-16):
get.docker.com reset the connection now and then (the third attempt got
through), and Docker's own mainland mirror of its apt repository served a
Packages index whose size did not match its Release file for over an hour
("Mirror sync in progress?"), so the install died in the Docker step both
times; the script's other mainland mirror lagged the package list the
script installs (no docker-model-plugin) and is no use as a second try.
So the fetch retries, and under --mirror cn the script is a first attempt:
when it fails, or cannot be fetched at all, the same five packages are
installed from USTC's mirror of the repository by hand, with its signing
key and an apt source written the way the script writes them. Off the
mainland an unreachable get.docker.com stays a plain refusal that says
what to do.
…he signing key's fingerprint pinned, not through the convenience script and its mirror fallbacks; node-gyp is pointed at the running Node's own headers; the node-role messages name the gateway machine's :80 The Docker step had grown three paths in a day (fetch get.docker.com with retries, run it against Docker's mainland mirror, fall back to USTC's mirror of the repository when that failed) around one unpinned remote script run as root — the one download in the installer without a checksum. Docker's own documentation gives production hosts the apt repository recipe; that is the one path now, with the repository's base URL the only thing --mirror cn changes, and the signing key's fingerprint compared with the one Docker publishes before apt is told to trust it (a mirror serves the key and the packages alike). Proven in an ubuntu:24.04 container against download.docker.com and against USTC's mirror. npm_config_nodedir was tied to the path this script unpacks its pinned Node into; the rule is simpler — a running Node whose prefix carries include/node has its own headers, and node-gyp is pointed there. Two --role node refusals still gave http://…:3677 as the example, the port ff0f372 established answers no other machine; the gateway unit's header comment said the same.
…tched from the branch the checkout tracks, not the tree's copy that belongs to the build being replaced An upgrade is "bring this machine to that commit", and only that commit's install.sh knows the host-side steps its build needs — a sysctl floor, a runtime flag, a new unit. apply() copied the tree's script, the one of the build being replaced, so twice in production the code arrived and the host-side step did not, until the new installer was re-run by hand (2026-09-01 --allow-suid, 2026-09-09 the inotify floor); the cut-over rules grew a standing "run install.sh by hand for a version that changed it". Now apply() fetches the tracked branch's head and writes its deploy/install.sh into the status directory before launching it — still outside the tree, whose file the script's own git pull replaces mid-run. The ref is the branch the checkout tracks (branch.<name>.remote and .merge, the way git pull reads it), for check() as well: it compared against origin main while install.sh pulled the checkout's upstream, and on a series-branch checkout the two disagreed — the version page read the build as ahead of main while the pull brought the branch. A detached HEAD or an untracked branch is an honest checkError, and one-click reports itself unavailable for that reason before a node could be told and fail. The mirror is judged from the remote's configured URL, before git's url.insteadOf rewrite, which lets the suite point the fixture's mirror URL at a local path and keep the fetch off the network.
…rsion it installs, so a host-side step the new version adds lands with the code
…its turn to upgrade comes — the script the updater now fetches, not the tree's
…e head of the branch the checkout tracks, which a detached HEAD on CI has none of
…n the one it was told on, not only when it reports the gateway's — a gateway upgraded again while the node built no longer reads that node as upgrading for twenty minutes and then stuck The row now records the commit the node ran when told (nodes.upgrade_told_build, migration 0006) beside the moment. A node back on any other commit is judged afresh: current, ahead, or behind and told again at its turn. Before, "still on the commit I was told on" was measured against the gateway's current build, so a fix pushed on the heels of the first push held the whole roll for twenty minutes with a reason that said the node had not moved when it had. A tell recorded before the column existed stands until the node reads current or ahead, as every tell once did.
…sion-cookie and admin-gate halves went to the gateway with the console in the second cut and were dead here, as was the @dormice/server/auth subpath nobody imported
What remains is tokensEqual and requireApiAuth(isCredential) — the one hook the node's two faces share. The gateway keeps its own auth.ts, the only copy of the console's credential code now; the stale comments in the daemon's copy ("the gateway never mounts the console") went with it.
…ounted in JavaScript — the check-in asks it every fifteen seconds, getHostMetrics on every poll, and a production ledger holds tens of thousands of rows countSandboxesByState(db) replaces countByState(listSandboxes(db)) at both call sites; the metrics sampler keeps loading the rows it samples one by one.
…, not the id the node had when the row was born — a node renamed since would otherwise list its older sandboxes under a node that no longer exists, a ghost beside the real one in a fleet's merged list Every row of a ledger is this node's; the row's node_id column stays as the birth record. The view takes the daemon's DORMICE_NODE_ID beside its endpoint (one Self for every row), and the shared schema says what the field means now.
… because an upgrade unit is already running on the machine — its previous upgrade still running doctor when the restarted daemon's first check-in is answered with the next tell — it tries again at the next check-in instead of dropping the tell with a warning Measured on the test machine 2026-09-16: an eight-second window between the daemon's first check-in on the new build and the end of the installer that brought it, and the gateway tells once — so a fix pushed on the heels of a fleet upgrade left the node reading upgrading for twenty minutes and then stuck, with the whole roll waiting behind it. Any other refusal (one-click unavailable, systemd-run failing) still drops the debt with the warning: trying again would change nothing there.
…nishing launches the next one at a later check-in
…— not two nullable fields read together: fleet.ts reads the row's two columns into `upgradeTold` or null, and rolling.ts's verdicts have no branch for a tell recorded without its commit (a row from before migration 0006 holding one reads as no tell); a tell that reads current or ahead while the node is still on the commit it was told on — the gateway went back to an older build — is forgotten too, so the node is not read stuck the day the gateway passes it again
…ts machine, and the gateway reads such a node as upgrading and does not tell it — instead of the node taking a tell it cannot launch as a debt and retrying: the eight seconds of doctor after the restarted daemon's first check-in are a fact the node knows (systemctl is-active dormice-upgrade, what its own 409 already asked), so the node states it and the gateway waits it out, telling at the first check-in that says the unit has ended; an installer run on a node by hand shows on the version page as upgrading for the same reason, and the check-in's owedUpgrade/waitingSaid memory goes
…units, not is-active: is-active makes systemd load the unit by name, and for a transient unit that has ended and been collected that reopens its fragment under /run/systemd/transient — gone — and logs a notice, two lines per ask (systemd 255, caught on the test machine the first quarter-hour the node reported `running` at every check-in: eleven thousand journal lines a day per node); list-units only lists what systemd holds in memory and loads nothing
…long that takes, never stuck: stuck is a tell with no unit behind it — the launch never happened, or the build failed and rolled back and the unit is gone — twenty minutes on, or a told node silent that long; read stuck while its installer still ran, a slow build moved the pointer and told the next node into the minute this one restarts in, with a remedy (put it back in line) that assumed the unit was gone. The twenty-minute clock was ever for the case where the gateway could not see the unit; since 04e1f4d the node says, and the gateway waits for what it says
… /console/envdToken, for good: the second is the name the verb was born with on the daemon, and clawsgo's server has minted its users' sandbox tokens through it under a Bearer token since July (22,000 calls a day on the Hong Kong machine) — the fleet gateway dropped the prefix (ee195db) and the first production cut-over answered that client 1,809 404s in its first thirty-five minutes (Hong Kong, 2026-09-17). A public path is a promise: renaming one earns nothing and costs every caller that kept its word. The node keeps its one name — ForwardOptions.path pins what the node is asked, a literal the route chose and not the caller's bytes
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft. The first of five steps toward a clustered Dormice: one gateway in front of N independent daemons. Not for merge on its own. The series lands together, and the production cut-over is a hand-run
install.sh, not the one-click updater.What this step adds
packages/gateway: the fleet's single door. It holds no sandbox state. Nodes check in every 15 s; a sandbox is found by asking every node in parallel (lookupSandbox); placement picks by (active + in-flight) / cores behind CPU, in-flight and data-disk gates. Two nodes holding one name answer 409 naming them; a silent node makes a new name answer 503 with Retry-After rather than risk a second copy.lookupSandboxverb; a check-in loop that runs only whenDORMICE_GATEWAY_ENDPOINTis set (unset means standalone, behaviour unchanged); the ledger lock handle held for the life of the process.e2e/src/gateway.test.ts: a black-box exam with a gateway in front of two fake-executor nodes.deploy/dormice-gateway.service, hand-installed for now.install.shroles come in a later step.Validated
pnpm testgreen in every package.The series
install.shroles and an image registry.Review commit by commit. Each message carries its reasoning, and the code comments carry the design.