Conversation
A container the durable object schedules has no image or size of its own: the wrangler containers block rejects instance_type under that policy, and the platform prepares images for the object to name rather than booting one itself. Every start() therefore has to carry an image, and carry a size if the workload needs one. This backend does that. The launch spec widens to the platform's own startup options, so entrypoint, labels and the snapshot fields reach start() too, and the caller names an image through `name`, which the host resolves against ctx.container.images. Consumers were previously reduced to patching ctx.container.start to inject any of it. Both launch paths build their spec through one helper. They reach the platform through different host methods, and a spec assembled separately in each place drifts silently: the restart drops what the initial start was given, and the replacement container comes up on a different image or at a different size than the one it replaced. The adoption digest grows the fields that cannot change on a running container, so a container launched at one size is no longer adopted by a caller asking for another. Snapshot fields stay out of it: they describe how a container was created rather than a property of the running one, and digesting them would relaunch a healthy container every time the stored handle moved on. An unresolvable image fails loudly rather than falling back. An empty images map means the container is platform-scheduled, which is a valid deployment this backend does not serve, so the error names the one that does.
🦋 Changeset detectedLatest commit: 5156b3a The changes in this PR will be included in the next version bump. This PR includes changesets to release 4 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
| res = await host.fetchPort(this.#options.containerPort, "http://container/connect", { | ||
| method: "POST", | ||
| headers: { | ||
| "content-type": "application/json", | ||
| authorization: `Bearer ${clientSecret}`, | ||
| }, | ||
| body: JSON.stringify({ | ||
| base: `http://${this.#options.egressHost}`, | ||
| health: EGRESS_HEALTH_PATH, | ||
| api: EGRESS_API_PATH, | ||
| healthTimeoutMs: remaining, | ||
| }), | ||
| }); |
There was a problem hiding this comment.
🔴 Connect request can hang forever
When /connect never responds, #postConnect ignores connectTimeoutMs and waits forever. Every workspace operation queued behind this connection remains blocked.
| res = await host.fetchPort(this.#options.containerPort, "http://container/connect", { | |
| method: "POST", | |
| headers: { | |
| "content-type": "application/json", | |
| authorization: `Bearer ${clientSecret}`, | |
| }, | |
| body: JSON.stringify({ | |
| base: `http://${this.#options.egressHost}`, | |
| health: EGRESS_HEALTH_PATH, | |
| api: EGRESS_API_PATH, | |
| healthTimeoutMs: remaining, | |
| }), | |
| }); | |
| res = await host.fetchPort(this.#options.containerPort, "http://container/connect", { | |
| method: "POST", | |
| headers: { | |
| "content-type": "application/json", | |
| authorization: `Bearer ${clientSecret}`, | |
| }, | |
| body: JSON.stringify({ | |
| base: `http://${this.#options.egressHost}`, | |
| health: EGRESS_HEALTH_PATH, | |
| api: EGRESS_API_PATH, | |
| healthTimeoutMs: remaining, | |
| }), | |
| signal: AbortSignal.timeout(remaining), | |
| }); |
Was this helpful? React with 👍 or 👎 to provide feedback.
| return { | ||
| enableInternet: spec.enableInternet, | ||
| envDigest: await digestEnv(spec.env), | ||
| ...(spec.instance === undefined ? {} : { instanceDigest: canonicalInstance(spec.instance) }), | ||
| ...(spec.name === undefined ? {} : { name: spec.name }), | ||
| }; |
There was a problem hiding this comment.
🔴 Redeployments retain stale container images
launchRecordFor treats a changed image behind the same name as the same launch. Image resolution occurs later, so redeployments adopt the old container and never run updated code.
Learn more
The adoption record compares requested launch data before the host resolves the selected image. A durable-object container can outlive a Worker redeployment, while the same image key can resolve to a new prepared image. The record also omits runtime-defining launch fields such as entrypoint, so those changes compare equal too. WorkspaceContainerAPI.start then adopts the surviving process instead of replacing it.
Example: Version 1 launches app as registry/app@sha256:old. Version 2 exposes app as registry/app@sha256:new. The persisted record still contains name: "app", so version 2 adopts the old process and continues running sha256:old.
Recommended fix: Build the adoption fingerprint from the effective launch configuration after resolving the image. Include the resolved image reference and every option that defines the running process, including entrypoint and timeout settings. Keep intentionally one-shot restore inputs outside the fingerprint only when adopting an already-running process preserves their completed effect.
Was this helpful? React with 👍 or 👎 to provide feedback.
| const stub = newWebSocketRpcSession( | ||
| ws as unknown as globalThis.WebSocket, | ||
| ) as RpcStub<WorkspaceRPC>; | ||
|
|
||
| // `closed` resolves on the first 'close' event from the underlying | ||
| // WebSocket. The Workspace listens for it and drops its cached | ||
| // handle so the next ready() call rebuilds against a fresh | ||
| // session. | ||
| let stopHeartbeat: (() => void) | undefined; | ||
| const closed = new Promise<void>((resolve) => { | ||
| let fired = false; | ||
| const onClose = () => { | ||
| if (fired) return; | ||
| fired = true; | ||
| stopHeartbeat?.(); | ||
| resolve(); | ||
| this.#handle = undefined; | ||
| }; | ||
| ws.addEventListener("close", onClose, { once: true }); | ||
| // Some runtimes fire 'error' without a follow-up 'close' on | ||
| // abrupt teardown; treat error as close too. | ||
| ws.addEventListener("error", onClose, { once: true }); | ||
| // capnweb's RPC layer can notice the session is broken (an | ||
| // abort frame, a malformed message) before the underlying | ||
| // WebSocket fires close. onRpcBroken closes that gap so the | ||
| // next ready() rebuilds against a fresh transport instead of | ||
| // waiting on a heartbeat or the next real RPC to discover | ||
| // the wedged session. | ||
| (stub as unknown as { onRpcBroken: (cb: (err: unknown) => void) => void }).onRpcBroken( | ||
| onClose, | ||
| ); | ||
| }); | ||
|
|
||
| if (this.#options.heartbeatIntervalMs > 0) { | ||
| stopHeartbeat = startHeartbeat({ | ||
| intervalMs: this.#options.heartbeatIntervalMs, | ||
| ping: () => (stub as unknown as WorkspaceRPC).sync.watermarks(), | ||
| onFailure: () => { | ||
| try { | ||
| ws.close(); | ||
| } catch { | ||
| // already closed; idempotent | ||
| } | ||
| }, | ||
| }); | ||
| } | ||
|
|
||
| const handle: BackendHandle = { | ||
| rpc: stub as unknown as WorkspaceRPC, | ||
| runtimeId, | ||
| closed, | ||
| close: async () => { | ||
| stopHeartbeat?.(); | ||
| // Dispose the root stub first. Per capnweb's docs, this is | ||
| // the documented way to shut a session down — it lets the | ||
| // RPC layer send a clean abort frame to the peer before | ||
| // the socket dies. Falling through to ws.close() is | ||
| // belt-and-braces for runtimes where the dispose path | ||
| // doesn't (yet) close the transport. | ||
| try { | ||
| (stub as unknown as Disposable)[Symbol.dispose]?.(); | ||
| } catch { | ||
| // already disposed; idempotent | ||
| } | ||
| try { | ||
| ws.close(); | ||
| } catch { | ||
| // already closed; idempotent | ||
| } | ||
| this.#handle = undefined; | ||
| }, | ||
| }; | ||
| this.#handle = handle; | ||
| return handle; |
commit: |
Add a runnable example for the durable-object-scheduled backend and use it in the primary setup documentation. The example shows how image and instance selection move from the Wrangler configuration into each container launch. Keep the platform-scheduled example under `container-legacy` for callers that still use that policy and for documentation tied to its exact setup.
3790b28 to
5156b3a
Compare
This PR introduces a new
ContainerRuntimebased on the new Cloudflare Container runtime.