Skip to content

Hard links copy the whole target node, symlinks included - #52

Merged
luke-lombardi merged 14 commits into
masterfrom
ll/hardlink-symlink
Sep 23, 2026
Merged

luke-lombardi merged 14 commits into
masterfrom
ll/hardlink-symlink

Conversation

@luke-lombardi

Copy link
Copy Markdown
Contributor

A hard link was replayed as a regular file with the target's attributes, so a hard link to a symlink kept the symlink mode but had no target and readlink failed with EINVAL. Nix's optimised store hard-links symlinks into /nix/store/.links, which broke every nixpacks-built image (unreadable libz.so.1 and friends). Regression test added.

Made with Cursor

luke-lombardi and others added 14 commits September 4, 2026 06:09
… content cache

Index-time seeding was removed in b25b736 so builds would not pay for
the extra disk write. Without it the first container to use a freshly
built layer materializes the whole layer before its first read returns
(10 s for a 2 GiB layer). The option lets the build path opt back in and
stores synchronously, since build workers exit right after indexing.

Co-authored-by: Cursor <cursoragent@cursor.com>
A sequential reader through FUSE issues 128 KiB reads, each of which was a
synchronous 1 MiB round trip to the content cache host; throughput was capped
by RTT. Windows are now 4 MiB and once a read crosses the second half of a
window the next one is fetched in the background, so the transfer overlaps
with the pages being consumed. Prefetches are singleflight-guarded and bounded
to four in flight per cache.

Co-authored-by: Cursor <cursoragent@cursor.com>
The store of a layer's decompressed tar into the content cache ran on the
indexing goroutine, holding one of the indexing slots for the upload's
duration (up to 10 s for a 500 MiB layer). Seeds now run through a
bounded seeder on the parent context, overlapping with the indexing of
the remaining layers, and are all awaited before the index is returned.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ontent cache

Range reads from the content cache go a window at a time and top out
well below the network; a mount that has read 64 MiB of a layer that way
(a shared library being mapped, a model being loaded) is going to read a
lot more of it. Stream the whole layer to the local disk cache in the
background at that point, so the rest of the reads are local. Layers
touched lightly stay in the cache.

Co-authored-by: Cursor <cursoragent@cursor.com>
One window ahead caps a sequential reader at two 4 MiB windows per round
trip, about 270 MiB/s on staging. Track where each layer's last window read
ended; a reader entering the following window is sequential and gets its
prefetch depth doubled, up to eight windows (32 MiB in flight), while a jump
elsewhere resets to one window. Raise the in-flight prefetch bound to 16 and
the window slots to 24 so the deeper prefetch is not evicted before use.

Co-authored-by: Cursor <cursoragent@cursor.com>
go-fuse sets max_read = MaxWrite (1 MiB) and splices fd-backed reads through
a pipe bounded by fs.pipe-max-size (1 MiB on stock hosts), so every 1 MiB
read from a local file view failed the size check, fell back to a copy and
logged 'trySplice: splice: want ... max pipe size'. Clamp MaxWrite to the
largest page-aligned size that fits.

Co-authored-by: Cursor <cursoragent@cursor.com>
A read that crossed a window boundary was given an oversized window key
(start, start+window+extra). That was a separate fetch of almost the same
bytes as the aligned window, and its odd end broke the sequential-stream
detection (lastEnd never matched the next window's start), so prefetch depth
reset to 1 on every straddling read. With 100 KiB files laid out back to back
every window was fetched three times and the reader never ramped up:
simulated 2000-file sequential read 275 ms -> 34 ms, 143 fetches -> 49.
Straddling reads now split across the aligned windows they cover.

Co-authored-by: Cursor <cursoragent@cursor.com>
A background warm streams a whole layer from the content cache to local
disk. On a host whose disk is the bottleneck those writes starve the
foreground reads that triggered the warm: a fresh sandbox importing torch
from a 5 GiB image took 17 s while its layers were being materialized,
against 2.5 s once they were local. The restore now records foreground
content-cache reads on the mount and holds itself to 64 MiB/s while one
happened in the last 250 ms, running flat out once the reader goes quiet.
Foreground materializations (a reader waiting on a whole layer) are never
paced.

Co-authored-by: Cursor <cursoragent@cursor.com>
The worker's layer prepare materializes every layer of an image eight at
a time two seconds after mount, which is the whole of a fresh sandbox's
first reads. Eight independently paced restores would still saturate a
400 MB/s disk, so the pacer's budget is now shared per mount, and Prepare
runs under the background-warm context so it is paced at all. A foreground
caller waiting on a layer (registry fallback joining an in-progress
restore) suspends pacing on the mount for as long as it waits.

Co-authored-by: Cursor <cursoragent@cursor.com>
Pacing every background restore to a fixed rate helped one sandbox on a
5 GiB image (17 s -> 5 s first import) but hurt ten sandboxes importing
torch on a fresh worker (2.4 s -> 6.4 s): the layer they all needed took
longer to become local, so their reads kept going to the network. The
pacer now exempts any layer the mount has read from in the active window;
those finish flat out and the shared budget applies only to the layers
nobody is touching, which is where the disk bandwidth was being wasted.

Co-authored-by: Cursor <cursoragent@cursor.com>
ca84433 exempted every layer the reader had touched in the active window
from the background pacer. A torch import on a fresh worker touches most
of a 14-layer, 5.4 GiB image inside the first second, so every restore
ran flat out and the import's own page reads queued behind 5 GiB of
background copy: first import 5-17 s on staging against 1.5 s once the
layers were local, and against Modal's 3-4.8 s for the same image.

Now only the most recently read layer (within the window) is exempt. The
layer the reader is in still becomes local at full speed, which is what
the ten-sandbox case needed, and the remaining restores share the
contended budget so the reads keep their bandwidth.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y waiting read

Materializing a layer streams it whole from the registry or the content
cache. One reset connection or throttled response failed the stream, and
the singleflight handed that error to every process whose read was waiting
on the layer: all of them got EIO at the same moment, which in a fresh
container is an ImportError. The read now retries the materialization up
to 4 times with backoff (waiters that shared the failed attempt retry too,
and one of them leads the next attempt); cancellation is not retried.

A read that does fail now leaves a warning with the path and cause, where
before the container saw EIO and the worker logged nothing.

Co-authored-by: Cursor <cursoragent@cursor.com>
A remote worker pulling from ECR gets one TCP flow per range, and from
the same site some S3 front-ends deliver 1-3 MiB/s while others deliver
60+. With 256 MiB parts and no supervision, one starved part held a 3 GiB
layer for 180 s; layers under 1 GiB were single-stream and hit the same
front-ends (a 400 MiB torch layer took 249 s).

- Threshold 1 GiB -> 32 MiB, parts 256 -> 32 MiB, 8 per layer, 24 total
  across layers on the worker.
- Straggler monitor: a part past a grace period that is below an absolute
  floor and far below the median of its peers/recent completions is
  cancelled and resumed from the bytes already written, on a new
  connection. In the tail (nothing left to hand out) any part well below
  the reference rate is resumed, since idle workers cost nothing.
- Front-ends that starved a range are avoided for 2 min by a custom
  dialer that spreads connections over the resolved addresses; pooled
  connections are dropped so keep-alive does not hand out the same one.
- Info summary per layer (duration, rate, aborts, retries); disk-headroom
  skip is now a warning.

On a skynet worker, 3 runs interleaved with single-stream:
  2.9 GiB model layer  17.3 / 14.6 / 15.1 s  vs  473 / 47 / 37 s
  383 MiB torch layer   4.9 /  5.7 /  3.8 s  vs  108 / 42 / 5.1 s

Co-authored-by: Cursor <cursoragent@cursor.com>
A hard link was replayed as a regular file with the target's attributes,
so a hard link to a symlink kept the symlink mode but had no target and
readlink failed with EINVAL. Nix's optimised store hard-links symlinks
into /nix/store/.links, which broke every nixpacks-built image.

Co-authored-by: Cursor <cursoragent@cursor.com>
@luke-lombardi
luke-lombardi merged commit 1c772eb into master Sep 23, 2026
1 check passed
luke-lombardi added a commit to beam-cloud/beta9 that referenced this pull request Sep 24, 2026
From the staging end-to-end run:

- `mcp install` writes the HTTP host the gateway reports (was the gRPC
host on hosted deployments).
- MCP `delete_app` stops deployments and containers (ran the REST
DeleteApp handler in-process); it only soft-deleted the app row.
- Deleting a database removes its app record; the dashboard kept a
stopped ghost app per deleted database.
- clip bumped to master for the hard-link-to-symlink fix
(beam-cloud/clip#52); nixpacks-built images had unreadable symlinks.
- runsc restore test: fake reports a pid only after the fake restore ran
(flaky on slow runners).

The container capability change that was here is now its own PR.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant