Skip to content

fix(vm): reclaim cancelled image preparation after worker exit - #4039

Open
shiju-nv wants to merge 1 commit into
fix/3950-vm-workload-identity-021400from
fix/3953-image-staging-cleanup-021400
Open

shiju-nv wants to merge 1 commit into
fix/3950-vm-workload-identity-021400from
fix/3953-image-staging-cleanup-021400

Conversation

@shiju-nv

@shiju-nv shiju-nv commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A timeout or cancellation can abort a sandbox's async preparation task while blocking filesystem work continues and temporary files remain. Run image and overlay preparation in an owned worker process. On cancellation, stop its process group and wait for cleanup before reclaiming files. Recover inactive attempts after driver restart.

Related Issue

Fixes #3953, accepted by a maintainer on 2026-09-30. Part of #3955. Depends on the workload-identity change in #4036.

Changes

  • Register each worker before yielding. Stop/delete join the cancelled task, terminate its process group, and wait for owned processes to exit. Inherited file locks protect files until descendants stop; uncertain cleanup retains staging and sandbox state.
  • Serialize shared cache publication and atomically publish complete disks, including writable-overlay retries. Reclaim marked inactive attempts at startup while preserving active attempts, committed caches and unmarked legacy staging.
  • Carry user/group selectors across the worker boundary and validate them against the persisted overlay owner before writes. Preserve successful worker status when stdout closes before exit, including Darwin's exited-process-group EPERM case.

Testing

  • mise run pre-commit passes locally. Focused local checks and the linked Branch Checks passed.
  • Unit tests added/updated
  • E2E tests added/updated (if applicable). Actual-worker integration tests passed; physical-VM testing has not been rerun for this candidate.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO). Both MG03 and MG06 signatures and DCO were verified.
  • Architecture docs updated (if applicable)

Run image preparation in an owned worker process, reserve its process
identity until cleanup completes, and protect staging with leases so
cancellation and recovery cannot race with another preparation attempt.

Signed-off-by: Shiju <shiju@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

@shiju-nv
shiju-nv changed the base branch from main to fix/3950-vm-workload-identity-021400 October 1, 2026 11:40
@shiju-nv
shiju-nv added this pull request to stack #4041 October 1, 2026 11:51
@shiju-nv
shiju-nv marked this pull request as ready for review October 1, 2026 19:20

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Reclaim inactive image staging after failed or cancelled MicroVM creation

1 participant