Skip to content

Mount the live virtio-fs share inside sandbox workloads - #222

Closed
Pedro Henrique Penna (ppenna) wants to merge 2 commits into
devfrom
fix/sandbox-virtiofs-mount
Closed

Pedro Henrique Penna (ppenna) wants to merge 2 commits into
devfrom
fix/sandbox-virtiofs-mount

Conversation

@ppenna

Copy link
Copy Markdown
Contributor

Fixes #216.

In sandbox mode (nvx_sandbox=1), OpenVMM attaches the --mount share and passes virtfs_dir, virtfs_tag, and virtfs_mode on the kernel command line, but nothing ever mounted it. The initramfs hook runs only outside sandbox mode, and nvx-init-agent ignored the tokens. The capability-stripped workload cannot mount the share itself. This PR mounts the share inside the workload root, and makes files that the workload creates owned by the host user who owns the export.

Commits

  1. Repin OpenVMM for caller-owned microVM virtio-fs requests (gitlink only): c9f659c36 → 65cc1c2a5, the head of virtiofs: perform microVM HostFs requests as the guest caller nanvix/openvmm#97 (open), which adds --mount-owner process|caller.
    • In caller mode, each FUSE request runs as the guest caller's UID and GID. The worker thread's supplementary groups and effective capabilities are dropped for the request.
    • Guest root is squashed to the export owner, and root-owned exports are rejected.
    • A request that cannot switch identity fails with EPERM.
    • process stays the default.
  2. Mount the live virtio-fs share inside sandbox workloads.

Guest agent (guest/common/nvx-init-agent)

  • After the overlay is assembled, and before identity resolution or any workload start, the agent mounts tag microvm at virtfs_dir inside $rootfs, using the requested mode plus nosuid,nodev. The mount lives below the overlay, so it survives --make-rprivate, unshare --mount, and chroot. The one-shot agent unmounts it before the overlay, and the managed agent inherits it for every exec.
  • The agent fails closed with NVX-SANDBOX-ERROR and status 125 on any of the following:
    • an unexpected tag or mode, or a non-canonical target;
    • /etc or /etc/machine-id, or a path under /proc, /sys, /dev, or /.nvx-agent, since the launcher manages or rewrites all of these;
    • an image symlink or non-directory in any component of the target (components are checked and created one at a time before anything else runs);
    • a failed mount.
  • Without --mount, behavior is unchanged.

CLI

  • nvx.py sandbox run|provision gains --mount, --mount-deny, and --mount-owner {process,caller}. The default is caller on Linux and process on Windows.
  • Managed state records the share with an absolute export root. Existing state directories still load.
  • Reserved targets are rejected before boot, and the share's bootstrap tokens count against the sandbox's 1024-byte command-line budget.
  • nvx.py run forwards an optional --mount-owner.

Tests

A new sandbox-filesystem microVM scenario boots the real sandbox agent over the Ubuntu EROFS layer with --mount /workspace,<dir>,rw --mount-deny <dir>/secrets and covers the issue's acceptance criteria:

  • a host-created file is visible to the workload;
  • files the workload writes, including a nested directory and a chmod, appear on the host while the VM is still running, with no copy-back;
  • a host edit made during the run is visible to the workload;
  • the denied path is absent;
  • the outcome report shows success with virtiofs_released: true;
  • an ro share rejects writes with EROFS;
  • the agent rejects a reserved target (/dev/...) and a target through the image's /bin symlink before any workload starts.

On Linux, the workload identity owns the export and the share uses caller, so the scenario also asserts the host owner of every file the workload creates. When run as root, it exports a dedicated 12345:12345 owner. The scenario formats its scratch image with mkfs.ext4 -d to define that identity, so runner provisioning now installs e2fsprogs. On Windows, it copies build/ubuntu-smoke-scratch.ext4.

Unit tests cover the launch contract, CLI forwarding and validation, state round-tripping, and the scenario's command construction and failure paths.

Answers to the questions in #216

  • Host UID/GID behavior is now documented in doc/run.md. By default OpenVMM performs requests as its own identity. With --mount-owner caller, workload-created files are owned by the workload identity; choosing the identity that owns the export, as AWF does with 1001, makes them owned by that host user. Caller mode requires OpenVMM to run as the export owner, or to hold CAP_SETUID and CAP_SETGID.
  • Credential-named files: a live share shows host files as-is, and NVX strips nothing. Use --mount-deny for secrets (option a).
  • Cache policy: zero entry and attribute lifetimes, plus direct I/O, so every lookup is a host round trip.
  • Egress: unchanged. The virtio-fs device adds no network path.
  • Multiple shares: the microVM ABI has exactly one fixed microvm slot, with a single access mode. A second share, such as a read-only tool cache next to a read-write workspace, needs a new ABI version. Documented in doc/design/current-limits.md.
  • Symlinks: guest symlink creation on the share still returns ENOTSUP. Link-creating tools such as npm ci must run in scratch. Left for a follow-up, and documented as a current limit.

Validation

Validated on baremetal hosts with the pinned OpenVMM 65cc1c2a5 and the NVX tree 222787f. The current head 450c548 differs from that tree only by the doc-only #220 change to doc/usage.md from the rebase onto dev.

Host test-openvmm-unit test-openvmm test-microvm (full, incl. sandbox-filesystem) CI Ubuntu-guest and sandbox smoke
prometheus32, Linux/KVM 4472 passed + doc tests 58 passed passed passed (including the KVM extended set)
prometheus30, Linux/MSHV (musl build) 4472 passed + doc tests 57 passed passed passed
prometheus28, Windows/WHP 4182 passed + doc tests 29 passed passed passed
  • Root path on Linux/KVM (sudo):
    • the lxutil credential tests and the virtiofs identity tests pass as root, including the loss of CAP_SETFCAP for a file that the switched identity owns;
    • sandbox-filesystem passes with a root OpenVMM switching every request to the 12345:12345 export owner.
  • Local checks: ruff check/format, pyright for Linux and Windows, all NVX unit suites, the pinned shellcheck/shfmt, and nvx.py verify. On the OpenVMM side, clippy -D warnings on Linux and Windows, rustdoc -D warnings, and xtask fmt.
  • Guest artifacts: the Alpine initramfs was rebuilt with Docker on prometheus32 and used on all three hosts.

A new v0.1.0-dev.* release will come from the dev push after merge.

Promote the `openvmm` submodule from `c9f659c36` to `65cc1c2a5`, the head
of nanvix/openvmm#97. That pull request adds `--mount-owner process|caller`
for the microVM `--mount` attachment:

- `lxutil: add scoped per-thread filesystem credentials` adds a guard that
  makes the calling thread perform filesystem operations exactly as another
  UID and GID. It switches the filesystem UID and GID, clears the
  supplementary groups, drops the effective capabilities, verifies every
  step, and restores the previous credentials when dropped.
- `virtiofs: perform microVM HostFs requests as the guest caller` runs each
  FUSE request under that guard when `caller` is selected. Guest root is
  squashed to the export owner, and a request that cannot switch identity
  fails with `EPERM`. `process` remains the default.

The sandbox agent and CLI changes that use the new option follow in the
next commit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 85d0240d-2f5e-4169-ae5d-540b7edc3b7b
In sandbox mode (`nvx_sandbox=1`), OpenVMM attaches the `--mount` share
and passes `virtfs_dir`, `virtfs_tag`, and `virtfs_mode` on the kernel
command line, but nothing ever mounts it. The initramfs hook runs only
outside sandbox mode, and `nvx-init-agent` ignores the tokens. The
capability-stripped workload cannot mount the share itself, so a host
directory exported with `--mount` never reaches the sandbox (#216).

Guest agent:

- After assembling the overlay, and before it resolves the workload
  identity or starts any workload, `nvx-init-agent` mounts the `microvm`
  tag at `virtfs_dir` inside the container root, with the requested mode
  plus `nosuid,nodev`. The mount lives below the overlay, so it survives
  the workload's private mount namespace and `chroot`. The one-shot agent
  unmounts it before the overlay, and the managed agent inherits it for
  every exec.
- The agent fails closed with `NVX-SANDBOX-ERROR` and status 125 if the
  tag or mode is unexpected, or if the target is not canonical. It also
  fails closed if the target is `/etc`, `/etc/machine-id`, or lies under
  `/proc`, `/sys`, `/dev`, or `/.nvx-agent`, all of which the container
  launcher manages or rewrites. The same applies when a component of the
  target in the image is a symbolic link or a non-directory, or when the
  mount fails. Components are checked and created one at a time before any
  workload exists, so an image symlink cannot redirect the mount outside
  the container root. Missing components are created as root-owned `0755`
  directories in scratch.
- Without `--mount`, sandbox behavior is unchanged.

CLI:

- `sandbox run` and `sandbox provision` accept `--mount`,
  `--mount-deny`, and `--mount-owner {process,caller}`. The defaults are
  `caller` on Linux, so workload-created files are owned by the export
  owner when the workload identity owns the export, and `process` on
  Windows. Managed state records the share with an absolute export root.
  Configurations provisioned before this change still load without a
  share.
- The launch contract rejects the reserved targets before boot, and
  counts the bootstrap tokens that OpenVMM appends against the sandbox's
  1024-byte command-line budget.
- `run` forwards an optional `--mount-owner`.

Tests:

- A new `sandbox-filesystem` microVM scenario boots the real sandbox agent
  over the Ubuntu EROFS layer. It shares `/workspace` read-write, hides
  `secrets/` with `--mount-deny`, and verifies what the issue asks for.
  The workload sees a host file. Its writes, nested directory, and `chmod`
  appear on the host while it still runs, with no copy-back. A host edit
  made during the run is visible to the workload. The denied path is
  absent. The outcome report records a clean virtio-fs teardown.
- On Linux the workload identity owns the export and the share uses
  `--mount-owner caller`, so the scenario also checks the host owner of
  every file the workload creates. When the tests run as root they export
  a dedicated owner, and the scenario formats its scratch image with an
  upper `/etc/passwd` that defines that identity.
- The scenario then verifies that a read-only share rejects workload
  writes with `EROFS`, and that the agent rejects a reserved target and a
  target that traverses the image's `/bin` symlink before any workload
  starts.
- Unit tests cover the launch contract, CLI forwarding and validation,
  managed-state round-tripping, and the scenario's command construction and
  failure handling.

Documentation:

- `doc/run.md` documents the share's caching, the fact that host files
  are shown as-is with `--mount-deny` for credentials, the unchanged
  egress policy, the single-share contract, `ENOTSUP` for guest symlink
  creation, host ownership with `--mount-owner`, and the sandbox mount
  behavior.
- `doc/design/current-limits.md`, the sandbox design, validation, and
  usage references follow.
- Runner provisioning installs `e2fsprogs` for the scenario's scratch
  image.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 85d0240d-2f5e-4169-ae5d-540b7edc3b7b

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It combines guest mount isolation with an unmerged security-sensitive OpenVMM credential-switching change whose exact-head checks remain in progress.

Review effort: Balanced
Findings: None

What changed in this PR

Enables live virtio-fs shares inside sandbox workloads, including caller-based host ownership, validation, lifecycle persistence, documentation, and cross-platform tests.

Changes:

  • Mounts validated shares inside the sandbox root with fail-closed behavior.
  • Adds CLI/state support for mount denial and ownership policy.
  • Adds end-to-end filesystem scenarios and required Linux tooling.
File Description
doc/​design/​current-limits.md Documents remaining virtio-fs limitations.
doc/​design/​sandbox-filesystem-and-agent-architecture.md Updates sandbox mount architecture.
doc/​design/​validation.md Documents live-share coverage.
doc/​run.md Explains ownership, security, and sandbox usage.
doc/​usage.md Adds CLI options and test prerequisites.
guest/​common/​init Clarifies sandbox-specific mounting.
guest/​common/​nvx-init-agent Validates, mounts, and unmounts sandbox shares.
openvmm Repins OpenVMM for caller-owned requests.
scripts/​nvx.py Adds mount options and forwarding.
scripts/​nvx_tools/​microvm_test_scripts/​sandbox-filesystem-read-only.sh.in Tests read-only enforcement.
scripts/​nvx_tools/​microvm_test_scripts/​sandbox-filesystem.sh.in Exercises live read-write sharing.
scripts/​nvx_tools/​microvm_tests.py Implements the end-to-end scenario.
scripts/​nvx_tools/​sandbox.py Defines and validates the mount contract.
scripts/​nvx_tools/​sandbox_lifecycle.py Persists mounts in managed state.
scripts/​setup/​README.md Documents the new tooling prerequisite.
scripts/​setup/​setup-linux-mshv.sh Installs e2fsprogs on MSHV hosts.
scripts/​setup/​setup-linux-runner.sh Installs e2fsprogs on Linux runners.
scripts/​test_microvm_tests.py Tests scenario orchestration and failures.
scripts/​test_nvx_tools.py Tests CLI, validation, and state handling.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Sandbox mode never mounts the live virtio-fs share (--mount) inside the workload container

2 participants