Skip to content

feat(supervisor): multi-tenant OpenShell proxy per workspace/namespace or broker #2385

Description

@dhirajsb

User Story

As an operator running many OpenShell sandboxes on Kubernetes, I want a multi-tenant OpenShell proxy service scoped to a workspace/namespace or broker, so that security enforcement can be operated and scaled independently of sandbox count, without allocating a dedicated supervisor/proxy pod for every sandbox.

Status: Draft RFC proposal for maintainer discussion. This updates the original shared-proxy proposal in this issue with the current two-pod Kubernetes architecture, explicit pod-capacity accounting, and workspace/namespace or broker scope. A numbered RFC remains subject to the OpenShell RFC process.

Problem Statement

The current Kubernetes deployment pairs each sandbox workload pod with a dedicated supervisor pod. The supervisor provides policy enforcement and proxying as well as trusted runtime functions. At scale, this couples the number of trusted service instances, their lifecycle, and their resource footprint to sandbox count, even for idle or low-traffic sandboxes.

The proposed change is to support one shared, replicated proxy service per workspace/namespace or per broker, replacing the mandatory 1:1 supervisor-pod-to-sandbox-pod relationship in shared mode. Per-sandbox policy, identity, credentials, audit attribution, and runtime isolation must remain distinct.

Merely forwarding dedicated supervisors through a shared upstream proxy does not achieve this goal: it leaves two pods per sandbox. Any retained supervision functionality must run as securely isolated per-sandbox sessions in shared trusted services, or another design that removes the dedicated supervisor pod while keeping trusted authority outside the agent workload.

Impact / Why This Matters

Two pods per sandbox halve pod-limited sandbox capacity

Worker nodes have a finite pod budget controlled by kubelet maxPods; available CPU and memory do not override that limit. Kubernetes documents maxPods as the maximum pods a kubelet can run, with a default of 110; actual cluster settings vary. See the Kubelet configuration reference.

Let:

  • P = available pod slots across eligible worker nodes after system pods, other workloads, and operational headroom are reserved.
  • N = concurrently resident sandbox workload pods, including idle/warm workload pods when present.
  • R = total shared proxy and any additional shared supervision replicas across the selected scopes, including replicas reserved for availability and rollout.
Model Pod consumption Pod-budget upper bound on sandbox capacity
Dedicated supervisor + workload 2N floor(P / 2)
Shared trusted service + workload N + R max(0, P - R)

The 1:1 model spends half the available pod slots on supervisor pods and therefore halves pod-limited sandbox capacity relative to one workload pod per sandbox. Sharing recovers most of those slots when R is small relative to N.

Illustrative arithmetic, not a benchmark: 10 workers configured for 110 pods each, with 10 slots per worker reserved for other pods and headroom, leave P = 1,000. The dedicated model supports at most 500 sandboxes. A shared deployment budgeted for 10 service replicas supports at most 990 sandboxes under the same pod budget. At 500 sandboxes, pod consumption falls from 1,000 to 510.

These are aggregate upper bounds. Per-node placement, CPU, memory, CNI/IP capacity, namespace quotas, storage, and service throughput can impose lower limits. Shared replicas consume pod slots too; many tiny workspace pools may save less than a broker pool. Raising maxPods requires validating node and networking capacity and preserves the two-pod multiplier.

Beyond pod slots, the 1:1 design repeats runtime memory, connection pools, policy caches, credential refresh activity, image startup, and scheduling/lifecycle work. Bursty creation and teardown amplify Kubernetes API and networking churn. A shared service can be sized for active traffic and availability instead of the total number of sandboxes.

Security operations should scale independently of untrusted workloads

A supervisor per sandbox multiplies trusted service instances and associated bootstrap material, policy state, certificate handling, and patch/rotation work. Short-lived sandbox churn also creates repeated opportunities for incomplete revocation or cleanup. These are operational security concerns, not evidence of a specific existing vulnerability.

A shared service enables centrally managed security updates, consistent enforcement, controlled credential handling, and independently hardened placement. Its rollout need not follow every workload's lifecycle. Agent pods must continue to lack provider credentials, proxy CA private keys, and gateway administrative authority.

Sharing increases the impact of a proxy compromise or isolation defect. The existing dedicated model has a valuable smaller process-level failure domain. Security improvement is therefore conditional on authenticated tenant separation, narrow credential authority, bounded resource use, and tested revocation. Workspace/namespace pools should be the first shared scope; broker-wide sharing requires explicit operator acceptance of its larger trust and failure domain. Dedicated placement remains available for stronger separation requirements.

Proposed Design

Operator workflow and placement scopes

Operators select placement through deployment configuration; sandbox callers cannot choose a pool or broaden its scope. Configuration names and APIs are to be agreed during design review.

Placement Intended scope Observable behavior
Dedicated One sandbox Existing topology and compatibility path
Workspace/namespace One explicitly bound workspace and its admitted namespace(s) Sandboxes share a replicated service within that trust boundary
Broker Sandboxes explicitly admitted by one broker, potentially across workspaces/namespaces A broker-scoped pool enforces separate identity, policy, credentials, and quotas for each tenant and sandbox

Here, broker means the operator's sandbox admission/lifecycle authority. An integration must bind it to concrete OpenShell gateway authority and immutable identities; a broker name or label is insufficient. Workspace and namespace are not assumed to be interchangeable: their mapping must be explicit, and a namespace containing multiple tenants still requires tenant isolation. Broker scope does not imply unrestricted cluster-wide trust.

  1. The operator provisions an available pool with scope, capacity, and security policy.
  2. Admission chooses the authorized pool and binds the sandbox runtime generation to it.
  3. The sandbox starts only after its protected channel, network boundary, policy, and scoped credentials are ready.
  4. Traffic is attributed to an authenticated sandbox session. Each sandbox's own rules apply regardless of which replica serves it.
  5. Operators independently scale, patch, rotate, and drain the pool. Sandbox deletion revokes its sessions and state without disrupting other sandboxes.

Security invariants

  • Authenticated identity: Every protected channel is mutually authenticated and bound to broker/gateway authority, tenant/workspace, immutable namespace and sandbox identity, runtime generation, and authorization epoch. IP addresses, namespace names, DNS, and workload-supplied headers cannot establish identity. Bootstrap identity must be short-lived or revocable and protected from agent access.
  • No cross-tenant selection: Policy, provider selection, credentials, middleware state, TLS material, caches, sessions, and audit records derive from trusted session context. A workload cannot select another sandbox's context through a header, route parameter, tunnel metadata, or connection reuse.
  • Least-privilege credential access: Credential resolution is constrained to the admitted sandbox, provider, destination/audience, and lifetime. Avoid a pool-wide credential granting unconstrained access to all tenant secrets. Never return provider secrets to workloads or logs. Partition connection pools whenever authentication, cookies, client certificates, or protocol state are identity-bound.
  • Preserve enforcement strength: Retain destination, TLS/L7, DNS/SSRF, middleware ordering, and process/binary identity checks. Process identity must come from trusted runtime mediation, not a workload assertion. A plain explicit HTTP proxy is insufficient if it drops existing guarantees.
  • No network bypass: Preserve the driver's default-deny workload fence and protected mediation channel. Do not introduce unrestricted direct egress as a fallback when the pool is unavailable. Evaluate IPv4/IPv6, DNS, alternate protocols, metadata endpoints, and same-cluster destinations. Do not assume workloads must initiate an outbound connection: the current runtime's channel direction and fence must be respected.
  • Fresh authorization: Policy snapshots and credential caches are versioned, bounded, and revoked on deletion, policy change, or identity rotation. Define a maximum revocation propagation interval and behavior for already-open streams. New admissions fail closed if authority is unavailable; existing sessions may continue only within explicitly valid authorization leases.
  • Protected service boundary: Run shared services outside agent-controlled pods, with narrow RBAC, restricted administration, hardened pod settings, and no unnecessary workload-creation authority. Namespace NetworkPolicy supplements application authorization; it cannot replace it.
  • Bounded sharing: Enforce per-sandbox and per-tenant limits on connections, bandwidth, buffers, queued work, policy evaluation, credential requests, and logs. Reserve capacity for control traffic so data-plane overload cannot block revocation or lifecycle actions.

Supervision and failover are part of the design

The current supervisor owns more than egress: it participates in runtime attachment, lifecycle/exec, and gateway relays. The driver records a supervisor pod UID, and the runtime pins a supervisor process instance. A Kubernetes Service alone cannot safely replace those bindings.

Shared placement must separate logical per-sandbox session ownership from physical replica identity while preserving generation fencing and exclusive ownership. A replacement replica must obtain fresh authority; it must not replay another replica's launch credentials. Concurrent owners, stale sessions after restart, and identity reuse after sandbox deletion must be rejected.

The RFC must decide which supervision functions share the pool and which remain in other trusted services. A design retaining one extra supervisor pod per sandbox is an intermediate migration step, not completion of the capacity objective. The workload-local runtime remains responsible for its existing isolation mechanisms; moving gateway or provider authority into the agent's trust boundary is out of scope.

Availability and scaling

Use multiple replicas distributed across failure domains, readiness based on usable policy/identity state, disruption budgets, and bounded connection draining. Scale on active connections, throughput, memory, queue depth, and latency; sandbox count alone is insufficient. Account for long-lived tunnels and skew when evaluating load balancing.

Distinguish transport reconnect from trusted session takeover. Replica loss must follow a defined fail-closed recovery path; whether a particular workload resumes or restarts depends on the approved ownership protocol. Do not promise transparent failover until it is demonstrated. Upstream operations must not be blindly replayed after an ambiguous connection failure.

Expose per-scope saturation, denials, policy revision, credential refresh errors, and revocation lag. Preserve tenant-authorized audit access and secret redaction while controlling metric cardinality.

Acceptance Criteria

  • Shared placement supports workspace/namespace and broker scopes with an explicit authority mapping; namespace scope can ship first.
  • For N sandboxes, shared-mode pod usage is N + R, with no dedicated supervisor/proxy pod per sandbox. Document any other shared infrastructure included in or outside R.
  • At a fixed worker count and maxPods, demonstrate increased resident sandbox capacity against the dedicated baseline, accounting for identical reserved headroom, warm pods, rollout replicas, and other resource constraints.
  • Two sandboxes with conflicting policies and different credentials retain their individual behavior while using the same replica.
  • Adversarial tests reject cross-tenant identity substitution, policy/cache confusion, credential leakage, connection-state reuse, stale generations, unauthorized pool access, and concurrent session ownership.
  • Existing runtime isolation, process-aware policy, DNS/SSRF, and protocol enforcement tests pass in shared mode. Direct-egress attempts remain blocked during normal operation and outages.
  • Deletion, policy revocation, and credential rotation take effect within documented limits, including established streams; unaffected tenants continue operating.
  • Replica failure, authority outage, network partition, rolling upgrades, and recovery demonstrate fail-closed behavior without stale-owner takeover.
  • Load tests include idle populations, connection bursts, long-lived sessions, and one abusive tenant. Report pod counts, CPU/memory, API churn, p95/p99 latency, connection limits, and the effect on other tenants. Set performance and recovery targets before rollout.
  • Dedicated mode remains supported, and a canary migration/rollback procedure preserves policy and identity boundaries. Rollback capacity planning reserves the extra pods required by the dedicated model.

Alternatives Considered

  • Keep dedicated proxies and reduce their resource requests: Preserves smaller failure domains but leaves the two-pod multiplier, instance count, and lifecycle coupling.
  • Increase maxPods or add workers: May extend capacity after infrastructure validation; incurs operational cost and leaves the structural overhead intact.
  • Put the trusted proxy in the workload pod: Saves a pod slot but changes network and credential isolation. Requires its own security analysis and does not provide an independently operated multi-tenant service.
  • Share only an upstream corporate proxy or credential refresh workers: Useful optimizations, but dedicated OpenShell supervisor pods remain.
  • Node-local proxy: Potential future placement with lower network distance; still needs authenticated tenant context and couples availability to node lifecycle.
  • Unrestricted cluster-wide pool: Broader authority and correlated failure than requested. Broker scope provides an explicit admission boundary; wider sharing should require separate review.

Open Questions and Rollout

  1. What is the precise mapping between broker, gateway, workspace, namespace UID, and tenant authority?
  2. Which current supervisor responsibilities move into shared services, and what is the session ownership/fencing protocol?
  3. How are bootstrap identities and TLS interception keys scoped, rotated, and revoked without widening credential authority?
  4. Which authorization caches and upstream connections can safely be shared, and which require sandbox or tenant partitioning?
  5. What revocation, availability, latency, and capacity targets gate broker-wide enablement?

Proposed sequence: agree on identity and ownership invariants; validate workspace/namespace pooling with adversarial and capacity tests; canary existing workloads with explicit rollback capacity; then enable broker scope only after cross-workspace isolation and noisy-neighbor testing pass. Retain dedicated placement for tenants requiring process-level separation.

Agent Investigation

Reviewed upstream main at commit 36b0386c92342569b0c87047b94aec4cf298470d on 2026-09-28:

  • Kubernetes driver: creates the supervisor pod separately from the workload pod and binds bootstrap state to both pod identities. This confirms the two-pod baseline and identifies admission/identity changes needed for pooling.
  • Sandbox architecture: describes protected runtime channels and supervisor instance pinning. Its standalone network-proxy role lacks several supervised-runtime capabilities and is not already a drop-in multi-tenant service.
  • Feature request template and RFC process: this issue supplies the design proposal for maintainer review before assignment of a formal RFC number.

This is a proposal supported by source inspection and capacity arithmetic. No implementation, cluster benchmark, or security validation is claimed.

Checklist

  • I've reviewed existing issues and the architecture docs
  • This is a design proposal, not a "please build this" request

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions