diff --git a/architecture/compute-runtimes.md b/architecture/compute-runtimes.md index ec6fff115a..5620cbd18f 100644 --- a/architecture/compute-runtimes.md +++ b/architecture/compute-runtimes.md @@ -357,6 +357,19 @@ Resource requirements enter the driver layer through `SandboxSpec.resource_requi can request a specific number of GPUs or the driver-specific default behaviour. For all in-tree drivers, this is equivalent to selecting a single GPU. +For Docker GPU sandboxes, the driver treats CDI specs as runtime metadata for +both outer injection and inner sandbox policy. It selects opaque CDI device IDs, +passes them to Docker, embeds a versioned CDI context in the protected workload +boundary configuration, and mounts daemon-reported CDI spec directories +read-only into the workload. The sandbox runtime resolves that context in the +workload mount namespace and derives Landlock paths and supplemental groups +from CDI `containerEdits` before launching agent processes. +Host-side CDI spec paths are diagnostic only and are never treated as sandbox +policy paths. +Kubernetes must not infer CDI device IDs from the `nvidia.com/gpu` resource +request; it needs a node-local selected-device handoff before using the same +workload-side resolver. + VM runtime state paths are derived only from driver-validated sandbox IDs matching `[A-Za-z0-9._-]{1,128}`. The gateway-owned VM driver socket uses a private `run/` directory plus Unix peer UID/PID checks. Standalone diff --git a/crates/openshell-core/src/cdi.rs b/crates/openshell-core/src/cdi.rs index cc89ada422..ba25715ebd 100644 --- a/crates/openshell-core/src/cdi.rs +++ b/crates/openshell-core/src/cdi.rs @@ -12,6 +12,12 @@ pub const CDI_CONTEXT_VERSION: u32 = 1; /// Base workload-runtime path under which compute drivers mount CDI specification directories. pub const CDI_SPEC_DIR_BASE: &str = "/run/openshell/boundary/cdi-specs"; +/// Return the workload-runtime path used for a CDI specification directory. +#[must_use] +pub fn cdi_spec_mount_path(index: usize) -> String { + format!("{CDI_SPEC_DIR_BASE}/{index}") +} + #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] pub struct CdiContext { pub version: u32, diff --git a/crates/openshell-driver-docker/README.md b/crates/openshell-driver-docker/README.md index a25c455eea..7720ef4722 100644 --- a/crates/openshell-driver-docker/README.md +++ b/crates/openshell-driver-docker/README.md @@ -85,6 +85,7 @@ LSM decisions remain authoritative. | Private named volumes | One carries the authenticated sandbox/supervisor channel. The other is mounted only into the supervisor and contains its JWT and private gateway credentials. | | In-memory `/run/openshell-supervisor-ca` tmpfs | Holds only the public supervisor CA certificate and trust bundle without making all of `/run` writable. | | CDI GPU request | Assigns the exact validated CDI devices requested by driver config or count-based selection. | +| CDI workload projection | For GPU sandboxes only, embeds selected device metadata in the protected boundary config and mounts Docker daemon CDI spec directories read-only into the workload. | ## Stop, Start, and Delete @@ -98,6 +99,20 @@ Delete force-removes both containers, the driver-owned runtime volumes, and the host-private runtime descriptor. Missing or altered descriptor and channel resources fail closed; the driver does not run an older combined-supervisor layout. +## CDI GPU Metadata + +Docker remains the source of truth for GPU injection. The driver selects opaque +CDI device IDs from `driver_config.cdi_devices` or the daemon's discovered CDI +inventory, then passes the same IDs to Docker with a CDI `DeviceRequest`. + +For a GPU sandbox, the driver writes those IDs and projected CDI spec directory +paths to a versioned context in the protected workload boundary configuration. +It bind-mounts the daemon-reported spec directories read-only into the workload. +The sandbox runtime resolves the selected devices in that mount namespace before +it launches agent processes and derives Landlock paths and supplemental groups +from CDI `containerEdits`. A missing context or spec directory fails workload +startup closed. Non-GPU sandboxes receive no CDI context or spec mounts. + ## Driver Config Mounts The gateway forwards the `docker` block from `--driver-config-json`. Supported diff --git a/crates/openshell-driver-docker/src/isolation.rs b/crates/openshell-driver-docker/src/isolation.rs index b5d4544e00..4150ca58f5 100644 --- a/crates/openshell-driver-docker/src/isolation.rs +++ b/crates/openshell-driver-docker/src/isolation.rs @@ -11,6 +11,7 @@ use std::collections::{BTreeMap, HashMap}; use std::net::IpAddr; use std::path::PathBuf; +use openshell_core::cdi::CdiContext; use openshell_isolation_interface::contract::{ BackendError, OuterFenceGuarantee, OuterFenceGuarantees, ResolvedWorkloadIdentity, }; @@ -70,6 +71,7 @@ pub struct DockerBoundarySpec { pub container_id: String, pub image_identity: String, pub gpu_requested: bool, + pub cdi_context: Option, pub listener_socket: PathBuf, pub control_socket: PathBuf, pub sandbox_tls: SandboxTlsServerConfig, @@ -118,7 +120,7 @@ impl DockerBoundarySpec { }, resource_claims: resource_claims.clone(), resource_claim_files: BTreeMap::new(), - cdi_context: None, + cdi_context: self.cdi_context, workload_identity: self.workload_identity.clone(), outer_fence: outer_fence.clone(), child_env: self.child_env, @@ -188,6 +190,13 @@ mod tests { container_id: "sha256:container".to_string(), image_identity: "sha256:image".to_string(), gpu_requested: true, + cdi_context: Some(CdiContext::new( + vec!["nvidia.com/gpu=0".to_string()], + vec![openshell_core::cdi::CdiSpecDirectory::new( + openshell_core::cdi::cdi_spec_mount_path(0), + "/var/run/cdi", + )], + )), listener_socket: PathBuf::from("/run/openshell/boundary/control.sock"), control_socket: PathBuf::from("/host/control.sock"), sandbox_tls: SandboxTlsServerConfig { @@ -224,6 +233,15 @@ mod tests { provisioned.runtime_descriptor.resource_claims[GPU_RESOURCE_CLAIM], "true" ); + assert_eq!( + provisioned + .boundary_config + .cdi_context + .as_ref() + .unwrap() + .selected_devices, + ["nvidia.com/gpu=0"] + ); assert_eq!( provisioned.boundary_config.outer_fence, provisioned.runtime_descriptor.outer_fence diff --git a/crates/openshell-driver-docker/src/lib.rs b/crates/openshell-driver-docker/src/lib.rs index 1f4c91f5e4..08482d0515 100644 --- a/crates/openshell-driver-docker/src/lib.rs +++ b/crates/openshell-driver-docker/src/lib.rs @@ -23,6 +23,7 @@ use bollard::query_parameters::{ }; use bytes::Bytes; use futures::{Stream, StreamExt}; +use openshell_core::cdi::{CdiContext, CdiSpecDirectory, cdi_spec_mount_path}; use openshell_core::config::DEFAULT_STOP_TIMEOUT_SECS; use openshell_core::driver_mounts; use openshell_core::driver_utils::{ @@ -311,12 +312,49 @@ struct DockerDriverRuntimeConfig { app_armor_profile: Option, } -#[derive(Debug, Clone, Copy)] +#[derive(Debug, Clone)] struct DockerGpuRuntimeCapabilities { cdi_supported: bool, + cdi_spec_dirs: Vec, wsl_all_gpu_fallback_enabled: bool, } +impl DockerGpuRuntimeCapabilities { + fn cdi_context(&self, gpu_device_ids: Option<&[String]>) -> Result, Status> { + let Some(gpu_device_ids) = gpu_device_ids.filter(|device_ids| !device_ids.is_empty()) + else { + return Ok(None); + }; + self.require_cdi_spec_dirs()?; + Ok(Some(CdiContext::new( + gpu_device_ids.to_vec(), + self.cdi_spec_dirs + .iter() + .enumerate() + .map(|(index, source)| CdiSpecDirectory::new(cdi_spec_mount_path(index), source)) + .collect(), + ))) + } + + fn cdi_spec_bind_strings(&self) -> Result, Status> { + self.require_cdi_spec_dirs()?; + Ok(self + .cdi_spec_dirs + .iter() + .enumerate() + .map(|(index, source)| format!("{source}:{}:ro,z", cdi_spec_mount_path(index))) + .collect()) + } + + fn require_cdi_spec_dirs(&self) -> Result<(), Status> { + if self.cdi_spec_dirs.is_empty() { + return Err(Status::failed_precondition( + "docker GPU sandboxes require Docker CDI spec directories reported by the daemon", + )); + } + Ok(()) + } +} #[derive(Clone)] pub struct DockerComputeDriver { docker: Arc, @@ -863,14 +901,13 @@ impl DockerComputeDriver { let info = docker.info().await.map_err(|err| { Error::execution(format!("failed to query Docker daemon info: {err}")) })?; - let cdi_supported = info - .cdi_spec_dirs - .as_ref() - .is_some_and(|dirs| !dirs.is_empty()); + let cdi_spec_dirs = info.cdi_spec_dirs.clone().unwrap_or_default(); + let cdi_supported = !cdi_spec_dirs.is_empty(); let cdi_gpu_inventory = docker_cdi_gpu_inventory(&info); let wsl_all_gpu_fallback_enabled = docker_info_reports_wsl2(&info); let gpu = DockerGpuRuntimeCapabilities { cdi_supported, + cdi_spec_dirs, wsl_all_gpu_fallback_enabled, }; validate_sandbox_pids_limit(docker_config.sandbox_pids_limit)?; @@ -941,7 +978,7 @@ impl DockerComputeDriver { supervisor_grpc_endpoint, ssh_socket_path: docker_config.ssh_socket_path.clone(), guest_tls, - gpu, + gpu: gpu.clone(), sandbox_pids_limit: docker_config.sandbox_pids_limit, enable_bind_mounts: docker_config.enable_bind_mounts, allow_driver_config: docker_config.allow_driver_config, @@ -1761,7 +1798,7 @@ impl DockerComputeDriver { &created.id, &image, &workload_identity, - gpu_devices.is_some(), + gpu_devices.as_deref(), ) .await { @@ -4521,7 +4558,7 @@ async fn prepare_docker_boundary_files( container_id: &str, image: &DockerImageMetadata, workload_identity: &ResolvedWorkloadIdentity, - gpu_requested: bool, + gpu_device_ids: Option<&[String]>, ) -> Result<(), Status> { let directory = docker_boundary_state_dir(sandbox, config)?; let workspace_root = driver_mounts::resolve_oci_workspace_root(&image.working_dir) @@ -4551,7 +4588,8 @@ async fn prepare_docker_boundary_files( verification_keys, container_id: container_id.to_string(), image_identity: image.id.clone(), - gpu_requested, + gpu_requested: gpu_device_ids.is_some_and(|device_ids| !device_ids.is_empty()), + cdi_context: config.gpu.cdi_context(gpu_device_ids)?, listener_socket: PathBuf::from(BOUNDARY_SOCKET_MOUNT_PATH), control_socket: PathBuf::from(BOUNDARY_SOCKET_MOUNT_PATH), sandbox_tls: SandboxTlsServerConfig { @@ -5694,7 +5732,10 @@ fn build_container_create_body_for_image( }), ..Default::default() }); - let user_bind_strings = docker_driver_bind_strings(driver_config)?; + let mut user_bind_strings = docker_driver_bind_strings(driver_config)?; + if gpu_device_ids.is_some_and(|device_ids| !device_ids.is_empty()) { + user_bind_strings.extend(config.gpu.cdi_spec_bind_strings()?); + } let device_requests = gpu_device_ids.map(|device_ids| { vec![DeviceRequest { driver: Some("cdi".to_string()), diff --git a/crates/openshell-driver-docker/src/tests.rs b/crates/openshell-driver-docker/src/tests.rs index c5036f4547..239d330cc1 100644 --- a/crates/openshell-driver-docker/src/tests.rs +++ b/crates/openshell-driver-docker/src/tests.rs @@ -23,6 +23,9 @@ use std::fs; use std::sync::Arc; use tempfile::TempDir; +const TEST_CDI_SPEC_DIR: &str = "/opt/openshell-test/cdi"; +const TEST_CDI_SPEC_DIR_ALT: &str = "/srv/openshell-test/cdi"; + fn test_launch_authentication() -> Vec { serde_json::to_vec(&SandboxLaunchAuthentication { supervisor: SupervisorAuthBundle { @@ -192,6 +195,7 @@ fn runtime_config() -> DockerDriverRuntimeConfig { }), gpu: DockerGpuRuntimeCapabilities { cdi_supported: false, + cdi_spec_dirs: Vec::new(), wsl_all_gpu_fallback_enabled: false, }, sandbox_pids_limit: openshell_core::config::default_sandbox_pids_limit(), @@ -202,6 +206,59 @@ fn runtime_config() -> DockerDriverRuntimeConfig { } } +fn runtime_config_with_cdi_spec_dirs(cdi_spec_dirs: &[&str]) -> DockerDriverRuntimeConfig { + let mut config = runtime_config(); + config.gpu.cdi_supported = !cdi_spec_dirs.is_empty(); + config.gpu.cdi_spec_dirs = cdi_spec_dirs + .iter() + .map(|path| (*path).to_string()) + .collect(); + config +} + +#[test] +fn gpu_runtime_builds_cdi_context_from_selected_devices_and_daemon_dirs() { + let config = runtime_config_with_cdi_spec_dirs(&[TEST_CDI_SPEC_DIR, TEST_CDI_SPEC_DIR_ALT]); + let selected_devices = vec![ + "nvidia.com/gpu=0".to_string(), + "nvidia.com/gpu=1".to_string(), + ]; + + let context = config + .gpu + .cdi_context(Some(&selected_devices)) + .unwrap() + .expect("GPU selection should produce a CDI context"); + + assert_eq!(context.selected_devices, selected_devices); + assert_eq!( + context.spec_dirs, + vec![ + CdiSpecDirectory::new(cdi_spec_mount_path(0), TEST_CDI_SPEC_DIR), + CdiSpecDirectory::new(cdi_spec_mount_path(1), TEST_CDI_SPEC_DIR_ALT), + ] + ); +} + +#[test] +fn gpu_runtime_omits_cdi_projection_without_selected_devices() { + let config = runtime_config_with_cdi_spec_dirs(&[TEST_CDI_SPEC_DIR]); + assert_eq!(config.gpu.cdi_context(None).unwrap(), None); +} + +#[test] +fn gpu_runtime_builds_read_only_cdi_bind_mounts() { + let config = runtime_config_with_cdi_spec_dirs(&[TEST_CDI_SPEC_DIR, TEST_CDI_SPEC_DIR_ALT]); + + assert_eq!( + config.gpu.cdi_spec_bind_strings().unwrap(), + vec![ + format!("{TEST_CDI_SPEC_DIR}:{}:ro,z", cdi_spec_mount_path(0)), + format!("{TEST_CDI_SPEC_DIR_ALT}:{}:ro,z", cdi_spec_mount_path(1)), + ] + ); +} + fn test_workload_identity() -> ResolvedWorkloadIdentity { ResolvedWorkloadIdentity::new( 1234, @@ -2453,8 +2510,7 @@ fn validate_sandbox_auth_accepts_launch_authentication() { #[test] fn build_container_create_body_maps_default_gpu_to_selected_cdi_device() { - let mut config = runtime_config(); - config.gpu.cdi_supported = true; + let config = runtime_config_with_cdi_spec_dirs(&[TEST_CDI_SPEC_DIR]); let mut sandbox = test_sandbox(); sandbox.spec.as_mut().unwrap().resource_requirements = Some(gpu_resources(None)); @@ -2479,6 +2535,15 @@ fn build_container_create_body_maps_default_gpu_to_selected_cdi_device() { request.device_ids.as_ref().unwrap(), &vec!["nvidia.com/gpu=1".to_string()] ); + let binds = create_body + .host_config + .as_ref() + .and_then(|host_config| host_config.binds.as_ref()) + .expect("GPU request should project CDI specs into the workload"); + assert!(binds.contains(&format!( + "{TEST_CDI_SPEC_DIR}:{}:ro,z", + cdi_spec_mount_path(0) + ))); } #[test] @@ -2501,8 +2566,7 @@ fn build_container_create_body_omits_devices_without_resolved_default_cdi_device #[test] fn build_container_create_body_passes_explicit_cdi_device_id_through() { - let mut config = runtime_config(); - config.gpu.cdi_supported = true; + let config = runtime_config_with_cdi_spec_dirs(&[TEST_CDI_SPEC_DIR]); let mut sandbox = test_sandbox(); let spec = sandbox.spec.as_mut().unwrap(); spec.resource_requirements = Some(gpu_resources(None)); @@ -2572,8 +2636,7 @@ fn build_container_create_body_rejects_empty_cdi_devices() { #[test] fn driver_default_gpu_selection_consumes_distinct_devices_for_creates() { - let mut config = runtime_config(); - config.gpu.cdi_supported = true; + let config = runtime_config_with_cdi_spec_dirs(&[TEST_CDI_SPEC_DIR]); let driver = test_driver_with_config(config); driver.gpu_selector.refresh( CdiGpuInventory::new(["nvidia.com/gpu=0", "nvidia.com/gpu=1"]), diff --git a/docs/how-it-works/sandboxes/runtimes.mdx b/docs/how-it-works/sandboxes/runtimes.mdx index cd3d576b52..fa347b67b0 100644 --- a/docs/how-it-works/sandboxes/runtimes.mdx +++ b/docs/how-it-works/sandboxes/runtimes.mdx @@ -91,6 +91,17 @@ Common options in `[openshell.drivers.docker]` are `socket_path`, `grpc_endpoint Docker Desktop must have host networking enabled, and it cannot use Enhanced Container Isolation. Set `grpc_endpoint` when sandboxes cannot reach the gateway on host loopback. For GPU sandboxes, configure Docker CDI before starting the gateway. +For Docker GPU/CDI sandboxes, OpenShell uses Docker's selected CDI device IDs +and daemon-reported CDI spec directories to build a workload CDI context. The +driver embeds the context in the protected boundary configuration and +bind-mounts the spec directories read-only into the workload. If context +creation fails, the driver removes the created state. If workload or supervisor +startup fails, it also removes the runtime resources before reporting the +failure. The sandbox runtime derives filesystem and supplemental group +requirements from CDI specs before launching agent processes. Non-GPU Docker +sandboxes do not receive the CDI context, spec mounts, or CDI-derived policy +changes. + ### Docker Mounts Mount existing named volumes or `tmpfs` through driver config. Label volumes so resource admission accepts them: