Skip to content

Measure what a sub-interpreter cell actually costs - #292

Merged
seanwevans merged 1 commit into
mainfrom
claude/festive-johnson-fd03cl-perf-honesty
Sep 17, 2026
Merged

seanwevans merged 1 commit into
mainfrom
claude/festive-johnson-fd03cl-perf-honesty

Conversation

@seanwevans

Copy link
Copy Markdown
Owner

First of a series of independent PRs working toward an honest multi-tenant Python execution fabric on free-threaded CPython. This one fixes the performance claims and adds the harness that produces them.

The problem

The performance snapshot quoted these against "the sub-interpreter backend":

Metric Value
Spawn latency 0.7 ms
Round-trip (1 kB) 70 µs
Max encrypted msgs/core 1.9 M/s
Baseline RSS 0.5 MiB

Two problems. The first two rows are numbers for the thread backend that backend="subinterpreter" selects today — not for a CPython sub-interpreter — so they understate the roadmap item that replaces it by more than an order of magnitude. The last two rows are not produced by anything in the repository.

What this adds

scripts/cell_cost.py measures the CPython primitives the planned sub-interpreter backend would be built on, so the roadmap can be argued from numbers taken on a real machine. Measured here on free-threaded CPython 3.14, 4-core container, --iterations 30:

Import surface in the cell create+exec (p50) close (p50) RSS/cell
bare (x = 1) 10.9 ms 3.9 ms 3.5 MiB
json, re, dataclasses 29.2 ms 7.8 ms 7.7 MiB
+ email, http.client, logging, argparse 56.8 ms 13.3 ms 13.3 MiB
Reference point p50
dispatch onto an already-warm pool cell 0.82 ms
fork() of a warm parent 1.62 ms

Three things follow, and they shape the rest of the series:

  • Creating a cell is not cheap. A sub-interpreter re-imports every module it uses with no copy-on-write sharing, so the import surface — not the interpreter object — dominates. A fork(), which is a real boundary, costs less than the cheapest possible cell.
  • Cells are only cheap when pooled. 0.82 ms is the number worth designing around, which argues for a pool of cells with pre-warmed import surfaces rather than an interpreter per request.
  • Parallelism comes from the build, not from the interpreters. Four cells scaled 3.9× — but four plain threads on that free-threaded build scaled 3.3× too. On a GIL build (3.13) those same threads scaled 0.96× while cells scaled 3.3×. So sub-interpreters buy parallelism on a GIL build, and buy a private sys.modules and a private set of globals per tenant on a free-threaded one.

A CPython bug worth knowing about

The lifecycle rows are gated to CPython 3.14+. Holding ~20 interpreters that have imported http.client or email.message and then destroying them aborts the process:

munmap_chunk(): invalid pointer

Reproducible on 3.13.12's private _interpreters; clean on 3.14's public concurrent.interpreters with the same workload. --allow-unstable-api measures anyway. This is a direct argument for targeting 3.14+ rather than the private module when the real backend lands.

Scope

Docs and a new script only — no runtime behaviour changes. The README's shipped-backend numbers are kept and relabelled as thread-backend figures; the two unreproducible rows are dropped rather than left standing.

Testing

  • tests/test_cell_cost.py — 6 new tests covering the reporting contract the README quotes from and the version guard (the guard is asserted via monkeypatch so it is checked on every Python in the matrix, not just 3.13).
  • Full suite: 536 passed, 14 skipped. The one failure in my container, test_apply_confinement_installs_seccomp_and_allows_normal_syscalls, is pre-existing and environmental — seccomp is unavailable here — and reproduces on unmodified main.
  • pre-commit run --all-files: isort, black, pylint, flake8, mypy all pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Fs1hTtmF4Hm9h617AG9Gse


Generated by Claude Code

The performance snapshot quoted 0.7 ms spawn and 0.5 MiB RSS against the
"sub-interpreter backend". Those are numbers for the thread backend that
`backend="subinterpreter"` selects today, not for a CPython sub-interpreter,
so they understate the roadmap item that replaces it by more than an order of
magnitude. Two of the four rows were also not produced by anything in the
repository.

Add scripts/cell_cost.py, which measures the primitives that backend would be
built on, and quote it. On free-threaded CPython 3.14 a cell costs 10.9 ms and
3.5 MiB bare, 56.8 ms and 13.3 MiB once it has imported a realistic module
surface -- against 1.62 ms for a fork(), which is a real boundary. The one
regime where a cell is cheap is dispatch onto an already-warm pooled
interpreter, at 0.82 ms, which is an argument for a cell pool rather than an
interpreter per request.

The script also separates two things that get conflated: on the free-threaded
build four plain threads scaled 3.3x against four cells' 3.9x, while on a GIL
build the same threads scaled 0.96x. Parallelism comes from the build;
sub-interpreters buy a private sys.modules and a private set of globals.

Lifecycle measurement is gated to CPython 3.14+. Holding ~20 interpreters that
imported http.client or email.message and then destroying them aborts the
process with `munmap_chunk(): invalid pointer` on 3.13's private
`_interpreters`; the same workload is clean on 3.14's public
`concurrent.interpreters`. `--allow-unstable-api` measures anyway.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fs1hTtmF4Hm9h617AG9Gse
@seanwevans
seanwevans merged commit 0234d1d into main Sep 17, 2026
9 of 19 checks passed
@seanwevans
seanwevans deleted the claude/festive-johnson-fd03cl-perf-honesty branch September 17, 2026 01:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants