Skip to content

Add new ContainerBackend based on new container runtime - #162

Open
aron-cf wants to merge 2 commits into
container-backend-legacyfrom
container-backend
Open

aron-cf wants to merge 2 commits into
container-backend-legacyfrom
container-backend

Conversation

@aron-cf

@aron-cf aron-cf commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

This PR introduces a new ContainerRuntime based on the new Cloudflare Container runtime.


Devin Review

A container the durable object schedules has no image or size of its
own: the wrangler containers block rejects instance_type under that
policy, and the platform prepares images for the object to name rather
than booting one itself. Every start() therefore has to carry an image,
and carry a size if the workload needs one.

This backend does that. The launch spec widens to the platform's own
startup options, so entrypoint, labels and the snapshot fields reach
start() too, and the caller names an image through `name`, which the
host resolves against ctx.container.images. Consumers were previously
reduced to patching ctx.container.start to inject any of it.

Both launch paths build their spec through one helper. They reach the
platform through different host methods, and a spec assembled
separately in each place drifts silently: the restart drops what the
initial start was given, and the replacement container comes up on a
different image or at a different size than the one it replaced.

The adoption digest grows the fields that cannot change on a running
container, so a container launched at one size is no longer adopted by
a caller asking for another. Snapshot fields stay out of it: they
describe how a container was created rather than a property of the
running one, and digesting them would relaunch a healthy container
every time the stored handle moved on.

An unresolvable image fails loudly rather than falling back. An empty
images map means the container is platform-scheduled, which is a valid
deployment this backend does not serve, so the error names the one
that does.
@aron-cf
aron-cf added this pull request to stack #163 September 23, 2026 15:51
@changeset-bot

changeset-bot Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 5156b3a

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 4 packages
Name Type
@cloudflare/computer Minor
@cloudflare/dofs Minor
@cloudflare/computer-rpc Minor
@cloudflare/computerd Minor

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 3 potential issues.

Devin Review

Comment on lines +680 to +692
res = await host.fetchPort(this.#options.containerPort, "http://container/connect", {
method: "POST",
headers: {
"content-type": "application/json",
authorization: `Bearer ${clientSecret}`,
},
body: JSON.stringify({
base: `http://${this.#options.egressHost}`,
health: EGRESS_HEALTH_PATH,
api: EGRESS_API_PATH,
healthTimeoutMs: remaining,
}),
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Connect request can hang forever

When /connect never responds, #postConnect ignores connectTimeoutMs and waits forever. Every workspace operation queued behind this connection remains blocked.

Suggested change
res = await host.fetchPort(this.#options.containerPort, "http://container/connect", {
method: "POST",
headers: {
"content-type": "application/json",
authorization: `Bearer ${clientSecret}`,
},
body: JSON.stringify({
base: `http://${this.#options.egressHost}`,
health: EGRESS_HEALTH_PATH,
api: EGRESS_API_PATH,
healthTimeoutMs: remaining,
}),
});
res = await host.fetchPort(this.#options.containerPort, "http://container/connect", {
method: "POST",
headers: {
"content-type": "application/json",
authorization: `Bearer ${clientSecret}`,
},
body: JSON.stringify({
base: `http://${this.#options.egressHost}`,
health: EGRESS_HEALTH_PATH,
api: EGRESS_API_PATH,
healthTimeoutMs: remaining,
}),
signal: AbortSignal.timeout(remaining),
});

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +62 to +67
return {
enableInternet: spec.enableInternet,
envDigest: await digestEnv(spec.env),
...(spec.instance === undefined ? {} : { instanceDigest: canonicalInstance(spec.instance) }),
...(spec.name === undefined ? {} : { name: spec.name }),
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Redeployments retain stale container images

launchRecordFor treats a changed image behind the same name as the same launch. Image resolution occurs later, so redeployments adopt the old container and never run updated code.

Learn more

The adoption record compares requested launch data before the host resolves the selected image. A durable-object container can outlive a Worker redeployment, while the same image key can resolve to a new prepared image. The record also omits runtime-defining launch fields such as entrypoint, so those changes compare equal too. WorkspaceContainerAPI.start then adopts the surviving process instead of replacing it.

Example: Version 1 launches app as registry/app@sha256:old. Version 2 exposes app as registry/app@sha256:new. The persisted record still contains name: "app", so version 2 adopts the old process and continues running sha256:old.

Recommended fix: Build the adoption fingerprint from the effective launch configuration after resolving the image. Include the resolved image reference and every option that defines the running process, including entrypoint and timeout settings. Keep intentionally one-shot restore inputs outside the fingerprint only when adopting an already-running process preserves their completed effect.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +330 to +403
const stub = newWebSocketRpcSession(
ws as unknown as globalThis.WebSocket,
) as RpcStub<WorkspaceRPC>;

// `closed` resolves on the first 'close' event from the underlying
// WebSocket. The Workspace listens for it and drops its cached
// handle so the next ready() call rebuilds against a fresh
// session.
let stopHeartbeat: (() => void) | undefined;
const closed = new Promise<void>((resolve) => {
let fired = false;
const onClose = () => {
if (fired) return;
fired = true;
stopHeartbeat?.();
resolve();
this.#handle = undefined;
};
ws.addEventListener("close", onClose, { once: true });
// Some runtimes fire 'error' without a follow-up 'close' on
// abrupt teardown; treat error as close too.
ws.addEventListener("error", onClose, { once: true });
// capnweb's RPC layer can notice the session is broken (an
// abort frame, a malformed message) before the underlying
// WebSocket fires close. onRpcBroken closes that gap so the
// next ready() rebuilds against a fresh transport instead of
// waiting on a heartbeat or the next real RPC to discover
// the wedged session.
(stub as unknown as { onRpcBroken: (cb: (err: unknown) => void) => void }).onRpcBroken(
onClose,
);
});

if (this.#options.heartbeatIntervalMs > 0) {
stopHeartbeat = startHeartbeat({
intervalMs: this.#options.heartbeatIntervalMs,
ping: () => (stub as unknown as WorkspaceRPC).sync.watermarks(),
onFailure: () => {
try {
ws.close();
} catch {
// already closed; idempotent
}
},
});
}

const handle: BackendHandle = {
rpc: stub as unknown as WorkspaceRPC,
runtimeId,
closed,
close: async () => {
stopHeartbeat?.();
// Dispose the root stub first. Per capnweb's docs, this is
// the documented way to shut a session down — it lets the
// RPC layer send a clean abort frame to the peer before
// the socket dies. Falling through to ws.close() is
// belt-and-braces for runtimes where the dispose path
// doesn't (yet) close the transport.
try {
(stub as unknown as Disposable)[Symbol.dispose]?.();
} catch {
// already disposed; idempotent
}
try {
ws.close();
} catch {
// already closed; idempotent
}
this.#handle = undefined;
},
};
this.#handle = handle;
return handle;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Run capnweb lifecycle soak coverage

The new backend duplicates long-lived capnweb session and disposal plumbing. Repository guidance calls for the stub soak harness when this lifecycle changes.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@pkg-pr-new

pkg-pr-new Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@cloudflare/computer@162

commit: 5156b3a

Add a runnable example for the durable-object-scheduled backend and use
it in the primary setup documentation. The example shows how image and
instance selection move from the Wrangler configuration into each
container launch.

Keep the platform-scheduled example under `container-legacy` for
callers that still use that policy and for documentation tied to its
exact setup.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant