Hard links copy the whole target node, symlinks included - #52
Merged
Merged
Conversation
… content cache Index-time seeding was removed in b25b736 so builds would not pay for the extra disk write. Without it the first container to use a freshly built layer materializes the whole layer before its first read returns (10 s for a 2 GiB layer). The option lets the build path opt back in and stores synchronously, since build workers exit right after indexing. Co-authored-by: Cursor <cursoragent@cursor.com>
A sequential reader through FUSE issues 128 KiB reads, each of which was a synchronous 1 MiB round trip to the content cache host; throughput was capped by RTT. Windows are now 4 MiB and once a read crosses the second half of a window the next one is fetched in the background, so the transfer overlaps with the pages being consumed. Prefetches are singleflight-guarded and bounded to four in flight per cache. Co-authored-by: Cursor <cursoragent@cursor.com>
The store of a layer's decompressed tar into the content cache ran on the indexing goroutine, holding one of the indexing slots for the upload's duration (up to 10 s for a 500 MiB layer). Seeds now run through a bounded seeder on the parent context, overlapping with the indexing of the remaining layers, and are all awaited before the index is returned. Co-authored-by: Cursor <cursoragent@cursor.com>
…ontent cache Range reads from the content cache go a window at a time and top out well below the network; a mount that has read 64 MiB of a layer that way (a shared library being mapped, a model being loaded) is going to read a lot more of it. Stream the whole layer to the local disk cache in the background at that point, so the rest of the reads are local. Layers touched lightly stay in the cache. Co-authored-by: Cursor <cursoragent@cursor.com>
One window ahead caps a sequential reader at two 4 MiB windows per round trip, about 270 MiB/s on staging. Track where each layer's last window read ended; a reader entering the following window is sequential and gets its prefetch depth doubled, up to eight windows (32 MiB in flight), while a jump elsewhere resets to one window. Raise the in-flight prefetch bound to 16 and the window slots to 24 so the deeper prefetch is not evicted before use. Co-authored-by: Cursor <cursoragent@cursor.com>
go-fuse sets max_read = MaxWrite (1 MiB) and splices fd-backed reads through a pipe bounded by fs.pipe-max-size (1 MiB on stock hosts), so every 1 MiB read from a local file view failed the size check, fell back to a copy and logged 'trySplice: splice: want ... max pipe size'. Clamp MaxWrite to the largest page-aligned size that fits. Co-authored-by: Cursor <cursoragent@cursor.com>
A read that crossed a window boundary was given an oversized window key (start, start+window+extra). That was a separate fetch of almost the same bytes as the aligned window, and its odd end broke the sequential-stream detection (lastEnd never matched the next window's start), so prefetch depth reset to 1 on every straddling read. With 100 KiB files laid out back to back every window was fetched three times and the reader never ramped up: simulated 2000-file sequential read 275 ms -> 34 ms, 143 fetches -> 49. Straddling reads now split across the aligned windows they cover. Co-authored-by: Cursor <cursoragent@cursor.com>
A background warm streams a whole layer from the content cache to local disk. On a host whose disk is the bottleneck those writes starve the foreground reads that triggered the warm: a fresh sandbox importing torch from a 5 GiB image took 17 s while its layers were being materialized, against 2.5 s once they were local. The restore now records foreground content-cache reads on the mount and holds itself to 64 MiB/s while one happened in the last 250 ms, running flat out once the reader goes quiet. Foreground materializations (a reader waiting on a whole layer) are never paced. Co-authored-by: Cursor <cursoragent@cursor.com>
The worker's layer prepare materializes every layer of an image eight at a time two seconds after mount, which is the whole of a fresh sandbox's first reads. Eight independently paced restores would still saturate a 400 MB/s disk, so the pacer's budget is now shared per mount, and Prepare runs under the background-warm context so it is paced at all. A foreground caller waiting on a layer (registry fallback joining an in-progress restore) suspends pacing on the mount for as long as it waits. Co-authored-by: Cursor <cursoragent@cursor.com>
Pacing every background restore to a fixed rate helped one sandbox on a 5 GiB image (17 s -> 5 s first import) but hurt ten sandboxes importing torch on a fresh worker (2.4 s -> 6.4 s): the layer they all needed took longer to become local, so their reads kept going to the network. The pacer now exempts any layer the mount has read from in the active window; those finish flat out and the shared budget applies only to the layers nobody is touching, which is where the disk bandwidth was being wasted. Co-authored-by: Cursor <cursoragent@cursor.com>
ca84433 exempted every layer the reader had touched in the active window from the background pacer. A torch import on a fresh worker touches most of a 14-layer, 5.4 GiB image inside the first second, so every restore ran flat out and the import's own page reads queued behind 5 GiB of background copy: first import 5-17 s on staging against 1.5 s once the layers were local, and against Modal's 3-4.8 s for the same image. Now only the most recently read layer (within the window) is exempt. The layer the reader is in still becomes local at full speed, which is what the ten-sandbox case needed, and the remaining restores share the contended budget so the reads keep their bandwidth. Co-authored-by: Cursor <cursoragent@cursor.com>
…y waiting read Materializing a layer streams it whole from the registry or the content cache. One reset connection or throttled response failed the stream, and the singleflight handed that error to every process whose read was waiting on the layer: all of them got EIO at the same moment, which in a fresh container is an ImportError. The read now retries the materialization up to 4 times with backoff (waiters that shared the failed attempt retry too, and one of them leads the next attempt); cancellation is not retried. A read that does fail now leaves a warning with the path and cause, where before the container saw EIO and the worker logged nothing. Co-authored-by: Cursor <cursoragent@cursor.com>
A remote worker pulling from ECR gets one TCP flow per range, and from the same site some S3 front-ends deliver 1-3 MiB/s while others deliver 60+. With 256 MiB parts and no supervision, one starved part held a 3 GiB layer for 180 s; layers under 1 GiB were single-stream and hit the same front-ends (a 400 MiB torch layer took 249 s). - Threshold 1 GiB -> 32 MiB, parts 256 -> 32 MiB, 8 per layer, 24 total across layers on the worker. - Straggler monitor: a part past a grace period that is below an absolute floor and far below the median of its peers/recent completions is cancelled and resumed from the bytes already written, on a new connection. In the tail (nothing left to hand out) any part well below the reference rate is resumed, since idle workers cost nothing. - Front-ends that starved a range are avoided for 2 min by a custom dialer that spreads connections over the resolved addresses; pooled connections are dropped so keep-alive does not hand out the same one. - Info summary per layer (duration, rate, aborts, retries); disk-headroom skip is now a warning. On a skynet worker, 3 runs interleaved with single-stream: 2.9 GiB model layer 17.3 / 14.6 / 15.1 s vs 473 / 47 / 37 s 383 MiB torch layer 4.9 / 5.7 / 3.8 s vs 108 / 42 / 5.1 s Co-authored-by: Cursor <cursoragent@cursor.com>
A hard link was replayed as a regular file with the target's attributes, so a hard link to a symlink kept the symlink mode but had no target and readlink failed with EINVAL. Nix's optimised store hard-links symlinks into /nix/store/.links, which broke every nixpacks-built image. Co-authored-by: Cursor <cursoragent@cursor.com>
luke-lombardi
added a commit
to beam-cloud/beta9
that referenced
this pull request
Sep 24, 2026
From the staging end-to-end run: - `mcp install` writes the HTTP host the gateway reports (was the gRPC host on hosted deployments). - MCP `delete_app` stops deployments and containers (ran the REST DeleteApp handler in-process); it only soft-deleted the app row. - Deleting a database removes its app record; the dashboard kept a stopped ghost app per deleted database. - clip bumped to master for the hard-link-to-symlink fix (beam-cloud/clip#52); nixpacks-built images had unreadable symlinks. - runsc restore test: fake reports a pid only after the fake restore ran (flaky on slow runners). The container capability change that was here is now its own PR. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A hard link was replayed as a regular file with the target's attributes, so a hard link to a symlink kept the symlink mode but had no target and readlink failed with EINVAL. Nix's optimised store hard-links symlinks into /nix/store/.links, which broke every nixpacks-built image (unreadable libz.so.1 and friends). Regression test added.
Made with Cursor