Skip to content

feat(sandbox): add per-sandbox idle stop and post-stop lifetime deletion timeouts #3804

Description

@dhirajsb

User Story

As an operator or user of OpenShell sandboxes on Kubernetes, I want to configure an idle timeout and a lifetime timeout per sandbox, so that unused sandboxes automatically stop and release compute resources while preserving attached PVCs, and sandboxes that remain stopped are eventually deleted without requiring an external reaper.

Problem Statement

Provide a two-stage, server-managed lifecycle policy:

  1. Idle timeout: stop a running sandbox after a configured interval without qualifying activity.
  2. Lifetime timeout: delete a sandbox after it has remained stopped for a configured interval.

For this request, lifetime timeout means retention after successful stop, measured from stopped_at; it is not a maximum age measured from creation and must not unexpectedly delete a running sandbox. API naming should make this distinction clear, for example stopped_lifetime_timeout or delete_after_stop if maintainers prefer.

Existing explicit stop/start/delete operations should remain the lifecycle authority. The requested feature automates those operations using persisted per-sandbox policy and deadlines.

Impact / Why This Matters

Idle sandboxes continue to occupy worker pod slots, requested CPU/memory, and potentially GPUs or other allocated devices. In the current Kubernetes topology, each running sandbox also has a dedicated supervisor pod, so an unused sandbox can retain two pod slots against worker maxPods capacity. See the capacity discussion in #2385 and the Kubernetes implementation in #3144, consolidated into #2942.

Stopping should reclaim runtime resources promptly without sacrificing the ability to resume from persistent storage. A separate post-stop lifetime gives users a recovery window while preventing abandoned sandbox records and sandbox-owned control resources from accumulating indefinitely.

Client-side timers fail when a client disconnects or crashes. External reapers duplicate authorization, activity tracking, race handling, and cleanup behavior across integrations. The policy should continue to work when no client is connected and across gateway restarts or failover.

Proposed Design

Per-sandbox configuration and workflow

Allow authorized users to set and inspect both durations through the sandbox API, with corresponding CLI/SDK support for creation and updates. The two settings are independent; omitted settings remain disabled for backward compatibility. Use explicit disable semantics and reject invalid or negative durations. Operators may define defaults and upper bounds without allowing sandbox users to exceed their authority.

Example workflow, using illustrative field names rather than prescribing final API syntax:

  • Create a sandbox with idle_timeout = 30m and lifetime_timeout = 24h.
  • If it becomes idle, stop it after 30 minutes and release its workload and dedicated supervisor pods, compute allocations, and associated runtime resources.
  • Keep the sandbox record, restart configuration, attached PVCs, and any minimal Kubernetes objects needed for a safe restart. Stopping must not merely leave resident pods that still consume pod slots.
  • After stop has completed, record stopped_at and calculate delete_at = stopped_at + 24h.
  • If the user starts the sandbox before deletion is committed, cancel that stopped interval's deletion deadline. The next completed stop begins a new retention interval.
  • If it remains stopped for 24 hours, invoke the normal sandbox deletion lifecycle and reconcile cleanup to completion.

A manual stop should start the same post-stop timer. Updating an enabled lifetime recalculates its deadline from the original stopped_at, rather than granting a fresh interval on every update. Enabling lifetime on an already-stopped sandbox must expose the resulting deadline; if reliable historical stop time is unavailable, persist and report the activation time as the initial anchor. No retroactive deletion is introduced for sandboxes without an enabled policy.

Define activity explicitly

The idle decision must reflect useful work, not merely the presence of a connection or a running container:

  • Accepted user exec/job operations and meaningful interactive or workload traffic count as activity.
  • An active managed exec/job must protect a legitimate long-running computation even when it produces no output or network traffic.
  • Health checks, readiness probes, gateway/supervisor heartbeats, telemetry, and list/get polling do not reset the timer. An open but inactive session alone does not keep a sandbox alive indefinitely.
  • A permanently running entrypoint or daemon is not by itself proof of useful activity. Document how background jobs express an authorized, bounded busy lease or equivalent signal when automatic observation is insufficient.
  • Activity timestamps are server-observed or come from authenticated trusted runtime reports, scoped to the sandbox generation. Workload-supplied timestamps cannot arbitrarily postpone cleanup.

Before rollout, document the exact qualifying signals and behavior when activity observations are unavailable. Missing telemetry must not be treated as proof of idleness. Start the initial idle interval once the sandbox is running/ready; provisioning deadlines remain separate.

Resource and storage semantics

Idle stop preserves attached PVCs and their data. Release workload and dedicated supervisor pods, their CPU/memory/device allocations, and runtime-only resources while retaining the minimum state needed to restart. Shared proxy pools and resources belonging to other sandboxes must remain untouched.

Lifetime expiry deletes the sandbox and its owned nonpersistent resources. It must not implicitly authorize deletion of externally managed or shared PVCs. Preserve attached PVCs by default for this automated policy; any deletion of sandbox-owned persistent storage requires an explicit, documented storage-retention setting. Ensure owner references and finalizers do not accidentally garbage-collect PVCs intended for retention. Report retained storage so operators can manage its separate lifecycle and cost.

Reliability, authorization, and races

Persist effective timeouts, last qualifying activity, stop completion time, deadlines, lifecycle generation, and cleanup progress. Reconciliation must be idempotent, bounded, and recover after restarts; it must use normal stop/delete authorization and lifecycle paths rather than deleting Kubernetes objects independently.

Re-check activity and the current lifecycle/policy version before committing an automatic stop. Serialize stop, start, timeout updates, and deletion so a stale expiry cannot delete a restarted sandbox or overwrite a newer policy. Once deletion is committed, reject a competing start with a clear outcome. Do not arm post-stop deletion from a stop request alone: wait for confirmed stop completion. Failed stops and partial deletion remain visible and retryable.

Expose effective policy, timestamps, pending actions, and reasons such as IdleTimeout and LifetimeTimeout through inspection/watch APIs and lifecycle events. Emit an audit trail for automatic actions and policy changes. Tenant-scoped authorization must apply to extending or disabling timers.

Acceptance Criteria

  • Two sandboxes can use different idle and lifetime durations; either timer can be disabled independently.
  • An inactive running sandbox automatically stops within a documented reconciliation tolerance; meaningful activity postpones idle stop, while heartbeats and management polling do not.
  • A legitimate long-running managed job is not stopped solely because it is silent; background-work behavior and bounded busy signaling are documented and tested.
  • Kubernetes workload and dedicated supervisor pods are removed on idle stop, releasing pod slots and compute/device allocations. Attached PVCs and data survive, and restarting the sandbox can reuse them.
  • Post-stop lifetime begins at confirmed stop completion, applies to manual and automatic stops, and deletes only a sandbox that remains stopped.
  • Starting before deletion commits cancels the old deadline; the next stop creates a new interval. Concurrent activity, start, policy update, and expiry cannot cause stale cleanup.
  • Lifetime deletion cleans up the sandbox record and owned nonpersistent resources without deleting retained, external, or shared PVCs or shared proxy resources.
  • Deadlines and retry progress survive gateway/controller restart and failover. Stop failures, partial deletion, and unavailable activity observations have documented, observable behavior.
  • API/CLI/SDK inspection shows effective durations, timestamps, deadlines, and reasons. Unauthorized users cannot change another sandbox's policy.
  • Existing sandboxes with neither timer enabled keep their current behavior. End-to-end Kubernetes tests verify actual pod removal, PVC data retention, restart, and eventual deletion.

Alternatives Considered

  • Absolute TTL from creation: Useful for a hard expiration lease, but does not provide activity-aware stopping and a recovery window after stop. See feat(sandbox): add server-managed expiration leases #2591.
  • External scheduled reaper: Duplicates lifecycle logic and requires additional permissions and consistent activity accounting.
  • Kubernetes TTL/direct pod deletion: Does not by itself implement OpenShell stop/start semantics, PVC retention, or gateway metadata cleanup.
  • Only idle stop: Reclaims compute but leaves abandoned stopped sandbox records indefinitely.
  • Immediate deletion on idle: Removes the recovery window and conflates compute reclamation with sandbox/storage retention.

Agent Investigation

Checklist

  • I've reviewed existing issues and the architecture docs
  • This is a design proposal, not a "please build this" request

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions