Skip to content

Automated Sync from main to stable - #17

Merged
rhods-devops-app[bot] merged 19 commits into
stablefrom
main
Oct 4, 2026
Merged

rhods-devops-app[bot] merged 19 commits into
stablefrom
main

Conversation

@rhods-devops-app

Copy link
Copy Markdown

Automated Sync from main to stable

This PR automatically syncs the main branch to the stable branch by opening a pull request from main into stable.

Sync Summary

  • Latest commit: def17b94 Merge remote-tracking branch 'upstream/main'

  • Total commits to sync: 19

  • Source: https://github.com/red-hat-data-services/OpenShell.git @ main

  • Target: https://github.com/red-hat-data-services/OpenShell.git @ stable

  • PR head: main

Commits to be synced

Merging

GitHub automerge is enabled for this pull request once required checks pass.

alangou and others added 19 commits October 2, 2026 21:04
Signed-off-by: Adrien Langou <alangou@nvidia.com>
…4105)

* fix(sandbox): replace lifetime exec cap with retry deadlines

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): remove exec recovery overview change

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(skills): remove exec recovery CLI skill change

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
)

* feat(cli): stream non-TTY exec input before EOF

Add --stream-stdin using the existing interactive exec RPC without a PTY. Preserve separate output streams and enforce the existing 4 MiB cumulative input cap while forwarding input.

Require explicit clean stdin EOF and drain the response through its final gRPC status. Cover held-open input, limits, cancellation, trailers, and default finite-input behavior with subprocess and live sandbox regressions.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(cli): distinguish exec cancellation from stdin EOF

Treat transport termination and response cancellation as separate test observations. Verify explicit stdin EOF through the shared frame writer and cover cancellation in the pinned Tonic decoder.

Signed-off-by: Shiju <shiju@nvidia.com>

* style(cli): use lazy optional stdin dispatch

Signed-off-by: Shiju <shiju@nvidia.com>

* test(cli): use imported duration in stdin EOF regression

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
Signed-off-by: red-hat-konflux <126015336+red-hat-konflux[bot]@users.noreply.github.com>
Co-authored-by: red-hat-konflux[bot] <126015336+red-hat-konflux[bot]@users.noreply.github.com>
Fixes NVIDIA#4040

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
The Docker driver passed a bare reference as CreateImageOptions.from_image
with no tag. The daemon interprets a tagless fromImage as a request for
every tag in the repository and pulls them all (issue NVIDIA#4029).

Normalize a pull reference by appending ':latest' when it has neither an
explicit tag nor a digest, matching 'docker pull' and the Podman driver.
Parsing inspects only the final path component so a registry port (e.g.
'registry:5000/team/app') is not mistaken for a tag and a digest-pinned
reference ('...@sha256:...') is left untouched. Applied at both pull
sites (pull_image and pull_runtime_image).

Signed-off-by: Udaya Tejas <udayatejas2004@gmail.com>
* fix(vm): enforce the configured workload identity

Reject conflicting policy users and groups before VM image preparation
and before guest attach or process startup changes state. Validate
supervisor policy updates against the protected VM workload identity.

Preserve the gateway CA transport and capability-free sandbox launcher.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(sandbox): clarify VM identity rejection fixtures

Name invalid user and group fixtures distinctly and move the final
workload identity into its group mismatch test.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): align VM identity startup with current APIs

Pass the optional rejection-log key for VM identity failures and keep
generic startup-write regressions free of VM identity constraints.

Repair the call sites after the branch rebase so the identity and cleanup
proposals compile against the current startup helpers.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(vm): restore inactive sandbox workload identity

Recover the persisted overlay owner before publishing stopped and terminal
sandboxes. Keep resources manageable when identity metadata is invalid.
Clarify fixed MicroVM ownership in policy-generation guidance.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(vm): flush identity fixture before restart

Persist the canonical identity file before readiness and report the observed
exec, canonical and file-owner identities before comparing them.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
Run image preparation in an owned worker process, reserve its process
identity until cleanup completes, and protect staging with leases so
cancellation and recovery cannot race with another preparation attempt.

Signed-off-by: Shiju <shiju@nvidia.com>
* feat(gateway): validate VM filesystem tools during preflight

Check required local VM tools through config preflight and share executable
resolution with VM image operations. Report selected paths and actionable
errors without creating gateway or sandbox state.

Bound probe output and execution time, and clean up probe descendants on
interruption. Preserve pure static validation and skip local tool checks
for remote driver endpoints and unrelated drivers.

Fixes NVIDIA#3951
Related to NVIDIA#3955

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(gateway): stabilize filesystem preflight checks

Combine identical filesystem-tool error arms and normalize rendered
diagnostics in command tests so terminal wrapping preserves assertions.
Describe driver TLS validation without depending on removed guest fields.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(gateway): serialize preflight fixture paths as TOML

Keep temporary paths quoted and escaped through the TOML serializer
instead of relying on Rust Debug formatting.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
…DIA#3772)

* fix(relay): close outbound relay stream when target closes first

Closes NVIDIA#3724

out_tx was cloned into the target-reading task, so the function's
own copy kept the outbound relay stream open until the client side
also ended. A target that closed a keep-alive connection never
reached the client as EOF, so reused connections hung forever.

Move out_tx into the target-reading task instead of cloning it, so
dropping it on target EOF ends the outbound stream right away.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* fix(relay): keep client uploads after target half-close

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
Signed-off-by: Thota shashank <thotashashank302@gmail.com>
* feat(server): separate image preparation and admission deadlines

Give sandbox image preparation its own deadline and start the admission
deadline after preparation finishes. Persist both phases across gateway
restarts and show recovery guidance for the phase that expired.

Preserve the current service authorization schema and regenerate the Go
bindings with the preparation timestamps.

Fixes NVIDIA#3952
Related to NVIDIA#3955

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): simplify preparation timeout fallback selection

Use lazy Option fallbacks while preserving timeout messages and retained
sandbox behavior.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(server): clarify admission timer prerequisites

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(compute): enforce deadlines during initial sandbox create

Release stalled create operations after preparation expires and preserve
the timeout diagnosis across late driver results. Keep failed-create cleanup
bound to its original attempt so another replica can retry safely.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(server): satisfy deadline regression lints

Drop the create-error mutex guard before matching its cloned value and use
idiomatic iteration and timeout matching in the deadline fixtures.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(server): retain ownership of pending provisioning operations

Keep submitted create and start requests alive after caller cancellation,
monitor failure, or preparation timeout. Persist request ownership before
dispatch and retain staged uploads while the driver response is pending.

Require timeout cleanup to stop compute after driver settlement without
discarding an active cleanup claim. Fence late result handling against newer
operations, preserve failed-start recovery, and defer automatic restart while
another request owns the sandbox. Expose pending ownership in CLI JSON.

Add ordered multi-replica and cancellation regressions for the review findings.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(openshell): preserve tracing and accept ready create responses

Carry the request span into the detached provisioning worker so compute
driver calls remain attached to their parent trace after task handoff.

Update the compensation regression to require no backend DELETE when the
durable cleanup claim fails, matching the operation ownership requirement.

Accept the gateway's current Ready snapshot when a sandbox becomes ready
before CREATE returns. Do not require the client to observe an earlier
provisioning phase. Cover the Ready-only watch and command attachment.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
…A#4034)

* fix(vm): validate launch credentials before preparation

Negotiate a launch-authentication requirement and reject incomplete gateway
configuration before driver validation, archive consumption or persistence.
Validate replacement VM credentials before stopping active compute and
prefer explicit signer configuration over local discovery.

Closes NVIDIA#3949. Part of NVIDIA#3955.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(vm): validate saved launch credentials before restore side effects

Signed-off-by: Shiju <shiju@nvidia.com>

* chore(gateway): satisfy launch preflight lint checks

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(gateway): complete the MicroVM launch-signing example

Include the gateway JWT paths in the standalone VM configuration and show
the output directory required by local certificate generation. Explain
the distinction between configuration preflight and launch-key validation.

Refs NVIDIA#3949. Part of NVIDIA#3955.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(vm): account for preparation state in launch credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
@rhods-devops-app
rhods-devops-app Bot enabled auto-merge October 4, 2026 01:46
@rhods-devops-app
rhods-devops-app Bot merged commit b802046 into stable Oct 4, 2026
25 of 33 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants