Conversation
Creating a sandbox from an image whose OCI USER resolves to root (for
example `USER root`, or a non-root user listed in group 0 via the image's
/etc/group) failed with an opaque, internal-sounding message:
IdentityResolutionFailed: descriptor error: workload identity must not
contain UID or GID zero
It named neither the image nor the offending component, leaked the internal
`descriptor error:` prefix, and offered no remediation.
The Docker and Podman identity resolvers now emit an actionable error before
the identity constructor rejects the zero value. The message names the image,
says which component is zero (user, primary group, or a supplementary group),
where the value came from (the image's USER, its /etc/passwd or /etc/group,
or the policy), and how to proceed (use a non-root image, or set
process.run_as_user and process.run_as_group). A new shared helper,
`root_identity_rejection_message`, keeps the wording consistent across both
drivers and is unit-tested without a Docker/Podman daemon.
This is a diagnostics and documentation change only. Root is still rejected
by `ResolvedWorkloadIdentity::new`, and an image with no USER and no policy
identity still resolves to 1000:1000; no UID/GID mapping, isolation, or trust
behavior is changed.
Also correct `runtimes.mdx`, which claimed images without a USER must set both
policy identity fields; it now matches `default-policy.mdx` and the code (no
USER falls back to UID/GID 1000, and a root USER is rejected unless the policy
sets a non-root identity).
Addresses the error-message and docs acceptance criteria of NVIDIA#4030. The CLI
"print the failure once" criterion is deferred pending maintainer
clarification (elezar) on whether "once" covers the final visible output or
every terminal write including transient spinner frames.
Signed-off-by: Udaya Tejas <udayatejas2004@gmail.com>
|
I have read the DCO document and I hereby sign the DCO. |
|
Hey @udsy19, thanks for the PR. I filed #4030 and ran this branch against it: gateway built from In the cases I ran, each new message (correctly) names the image, which part is zero, and where it came from, and A few things about the advice that follows
Every one of the new messages ends with that same sentence. On the netshoot image from the issue (UID 0) it fits: 1. Supplementary group 0: following the advice can give the same message back. FROM nvcr.io/nvidia/base/ubuntu:24.04
RUN useradd --uid 1001 --user-group --no-create-home app && usermod --append --groups root app
USER appI set both fields to the image's own
|
| Image | Both fields set, as advised | Only run_as_group set |
|---|---|---|
USER 1000:0 |
1000/1000: starts, uid=1000(ubuntu) gid=1000(ubuntu) |
1001: starts, uid=1000(ubuntu) gid=1001 |
USER app, where app has GID 0 in /etc/passwd |
1001/1001: starts, uid=1001(app) gid=1001 |
1001: starts, uid=1001(app) gid=1001 |
So nothing is broken here. For these two images the message just asks for more than is needed. Maybe it could ask for process.run_as_group only.
3. "Use an image with a non-root USER" also shows up for images whose USER is already non-root. That's the image in 1 and the /etc/passwd image in 2, both USER app.
Maybe the advice could depend on which part is zero, the way the first half of the message already does.
All the policy files are the one above with only the process section changed.
Summary
Creating a sandbox from an image whose effective user/group resolves to 0 (e.g. a
USER rootimage, or root via/etc/passwd//etc/group) fails with an opaquedescriptor error: ... UID or GID zeromessage that doesn't say what's wrong or how to fix it. OpenShell correctly refuses to run a workload as root — this PR keeps that rejection but makes the error actionable, and fixes a docs table that disagreed with the code.Fixes #4030.
Scope (per @elezar's triage)
This is a diagnostics + CLI-presentation + docs change. Root rejection stays enforced; no UID/GID resolution, mapping, or isolation behavior changes.
What changed
root_identity_rejection_message+IdentityComponentOriginenum inopenshell-isolation-interface(contract.rs). It only explains an already-rejected identity — it does not touchResolvedWorkloadIdentity::new, which remains the enforced invariant. The message names the image, says which component is zero (user / primary group / supplementary group), and where its value came from (imageUSER,/etc/passwd,/etc/group, or policy), plus remediation (process.run_as_user/run_as_group); the supplementary-group-0 case notes that overridingrun_as_groupalone doesn't remove the membership.driver-docker/src/lib.rs,driver-podman/src/isolation.rs) compute the origin, pull the image reference, and return the actionable error directly — so the internaldescriptor error:prefix no longer leaks. Identity computation is byte-for-byte unchanged; no-USER-no-policy still resolves to1000:1000.docs/how-it-works/sandboxes/runtimes.mdxidentity table corrected to matchdefault-policy.mdxand the code (no-USER →1000; rootUSERrejected unless policy overrides).Deferred
@elezar flagged the "print the diagnostic once (TTY)" criterion as ambiguous ("Human Decision Required"). I've left the CLI print-dedup out of this PR pending your clarification on whether "once" covers only final visible output or every terminal write (e.g. transient spinner frames) — happy to add it in a follow-up once the intended semantics are settled.
Tests
openshell-isolation-interface: 5 helper tests (image-USER root, policy root, primary-GID-0 from passwd, supplementary-group-0, non-root accepted).openshell-driver-docker: +2 (USER root, supplementary-group-0) and the existing root-rejection assertion updated.openshell-driver-podman: strengthened to assert the actionable message.All pass;
cargo fmt --checkandcargo clippyclean on the three crates.