Skip to content

fix(compute): explain the root-identity rejection instead of an opaque UID/GID error - #4170

Open
udsy19 wants to merge 1 commit into
NVIDIA:mainfrom
udsy19:fix/clear-error-user-root-image
Open

udsy19 wants to merge 1 commit into
NVIDIA:mainfrom
udsy19:fix/clear-error-user-root-image

Conversation

@udsy19

@udsy19 udsy19 commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Summary

Creating a sandbox from an image whose effective user/group resolves to 0 (e.g. a USER root image, or root via /etc/passwd//etc/group) fails with an opaque descriptor error: ... UID or GID zero message that doesn't say what's wrong or how to fix it. OpenShell correctly refuses to run a workload as root — this PR keeps that rejection but makes the error actionable, and fixes a docs table that disagreed with the code.

Fixes #4030.

Scope (per @elezar's triage)

This is a diagnostics + CLI-presentation + docs change. Root rejection stays enforced; no UID/GID resolution, mapping, or isolation behavior changes.

What changed

  • New pure helper root_identity_rejection_message + IdentityComponentOrigin enum in openshell-isolation-interface (contract.rs). It only explains an already-rejected identity — it does not touch ResolvedWorkloadIdentity::new, which remains the enforced invariant. The message names the image, says which component is zero (user / primary group / supplementary group), and where its value came from (image USER, /etc/passwd, /etc/group, or policy), plus remediation (process.run_as_user / run_as_group); the supplementary-group-0 case notes that overriding run_as_group alone doesn't remove the membership.
  • The Docker and Podman resolvers (driver-docker/src/lib.rs, driver-podman/src/isolation.rs) compute the origin, pull the image reference, and return the actionable error directly — so the internal descriptor error: prefix no longer leaks. Identity computation is byte-for-byte unchanged; no-USER-no-policy still resolves to 1000:1000.
  • docs/how-it-works/sandboxes/runtimes.mdx identity table corrected to match default-policy.mdx and the code (no-USER → 1000; root USER rejected unless policy overrides).

Deferred

@elezar flagged the "print the diagnostic once (TTY)" criterion as ambiguous ("Human Decision Required"). I've left the CLI print-dedup out of this PR pending your clarification on whether "once" covers only final visible output or every terminal write (e.g. transient spinner frames) — happy to add it in a follow-up once the intended semantics are settled.

Tests

  • openshell-isolation-interface: 5 helper tests (image-USER root, policy root, primary-GID-0 from passwd, supplementary-group-0, non-root accepted).
  • openshell-driver-docker: +2 (USER root, supplementary-group-0) and the existing root-rejection assertion updated.
  • openshell-driver-podman: strengthened to assert the actionable message.
    All pass; cargo fmt --check and cargo clippy clean on the three crates.

Creating a sandbox from an image whose OCI USER resolves to root (for
example `USER root`, or a non-root user listed in group 0 via the image's
/etc/group) failed with an opaque, internal-sounding message:

    IdentityResolutionFailed: descriptor error: workload identity must not
    contain UID or GID zero

It named neither the image nor the offending component, leaked the internal
`descriptor error:` prefix, and offered no remediation.

The Docker and Podman identity resolvers now emit an actionable error before
the identity constructor rejects the zero value. The message names the image,
says which component is zero (user, primary group, or a supplementary group),
where the value came from (the image's USER, its /etc/passwd or /etc/group,
or the policy), and how to proceed (use a non-root image, or set
process.run_as_user and process.run_as_group). A new shared helper,
`root_identity_rejection_message`, keeps the wording consistent across both
drivers and is unit-tested without a Docker/Podman daemon.

This is a diagnostics and documentation change only. Root is still rejected
by `ResolvedWorkloadIdentity::new`, and an image with no USER and no policy
identity still resolves to 1000:1000; no UID/GID mapping, isolation, or trust
behavior is changed.

Also correct `runtimes.mdx`, which claimed images without a USER must set both
policy identity fields; it now matches `default-policy.mdx` and the code (no
USER falls back to UID/GID 1000, and a root USER is rejected unless the policy
sets a non-root identity).

Addresses the error-message and docs acceptance criteria of NVIDIA#4030. The CLI
"print the failure once" criterion is deferred pending maintainer
clarification (elezar) on whether "once" covers the final visible output or
every terminal write including transient spinner frames.

Signed-off-by: Udaya Tejas <udayatejas2004@gmail.com>
@udsy19
udsy19 requested review from a team, derekwaynecarr, mrunalp and sjenning as code owners October 3, 2026 20:53
@copy-pr-bot

copy-pr-bot Bot commented Oct 3, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@udsy19

udsy19 commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

I have read the DCO document and I hereby sign the DCO.

@mpereira

mpereira commented Oct 4, 2026 •

Copy link
Copy Markdown

Hey @udsy19, thanks for the PR.

I filed #4030 and ran this branch against it: gateway built from main at 71c3cd957 with this diff applied (head 0e24fe61b), Docker driver, macOS 26.2 arm64, Docker Engine 29.4.0.

In the cases I ran, each new message (correctly) names the image, which part is zero, and where it came from, and descriptor error is gone.

A few things about the advice that follows

Use an image with a non-root USER, or set process.run_as_user and process.run_as_group ...

Every one of the new messages ends with that same sentence. On the netshoot image from the issue (UID 0) it fits: run_as_user: "1000" alone fails and asks for run_as_group. For the GID 0 cases it doesn't quite fit. There's three separate things, and I'd say the first one is the important one:

1. Supplementary group 0: following the advice can give the same message back.

FROM nvcr.io/nvidia/base/ubuntu:24.04
RUN useradd --uid 1001 --user-group --no-create-home app && usermod --append --groups root app
USER app

I set both fields to the image's own app user, which has a non-root UID and GID:

policies/full-1001.yaml: the default policy plus run_as_user: "1001" and run_as_group: "1001"
version: 1
filesystem_policy:
  include_workdir: true
  read_only:
    - /bin
    - /usr
    - /lib
    - /proc
    - /dev/urandom
    - /etc
    - /var/log
  read_write:
    - /tmp
    - /dev/null
landlock:
  compatibility: best_effort
process:
  run_as_user: "1001"
  run_as_group: "1001"

The same message comes back, asking me to set the two fields I had just set:

$ openshell sandbox create --name supp-same-full-f --detach --from root-group-test:supplementary --policy policies/full-1001.yaml

Created sandbox: supp-same-full-f

  [0.0s] Requesting compute...
[0.0s] Sandbox allocated
[0.0s] Pulling image root-group-test:supplementary
[0.0s] Image already present
Error:   × sandbox entered error phase while provisioning: IdentityResolutionFailed:
  │ image 'root-group-test:supplementary' lists the workload user in group 0
  │ (root) through its `/etc/group`. Overriding `process.run_as_group` alone
  │ does not remove this supplementary membership. OpenShell never runs a
  │ workload as root. Use an image with a non-root `USER`, or set
  │ `process.run_as_user` and `process.run_as_group` to a non-root identity in
  │ the creation policy.

That's because the image's /etc/group lists app as a member of group 0 (root:x:0:app), and that comes along however app is selected.

What gets past it is selecting a different user, one that isn't a member of group 0. The base image's ubuntu user (UID 1000) isn't, so I changed the process section to just:

process:
  run_as_user: "1000"

That starts, and id in the sandbox shows no group 0:

$ openshell sandbox exec supp-user-1000-f -- id
uid=1000(ubuntu) gid=1000(ubuntu) groups=1000(ubuntu),4(adm),20(dialout),24(cdrom),25(floppy),27(sudo),29(audio),30(dip),44(video),46(plugdev)

So for this case, maybe the message could say that instead of the shared sentence: pick a user that isn't a member of group 0.

2. Primary GID 0: the advice works, but one field was enough.

This is a different case from 1. Here group 0 is the user's primary group, and the message is the "resolves to a root primary group ... selects GID 0" one, followed by the same advice. Two images:

Image Both fields set, as advised Only run_as_group set
USER 1000:0 1000/1000: starts, uid=1000(ubuntu) gid=1000(ubuntu) 1001: starts, uid=1000(ubuntu) gid=1001
USER app, where app has GID 0 in /etc/passwd 1001/1001: starts, uid=1001(app) gid=1001 1001: starts, uid=1001(app) gid=1001

So nothing is broken here. For these two images the message just asks for more than is needed. Maybe it could ask for process.run_as_group only.

3. "Use an image with a non-root USER" also shows up for images whose USER is already non-root. That's the image in 1 and the /etc/passwd image in 2, both USER app.

Maybe the advice could depend on which part is zero, the way the first half of the message already does.

All the policy files are the one above with only the process section changed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: sandbox from a USER root image fails with an unexplained "UID or GID zero" error

2 participants