Skip to content

websocket: Speed up masking (pure-Python and C) - #3774

Open
bdarnell wants to merge 2 commits into
tornadoweb:masterfrom
bdarnell:claude/quirky-lamport-0u7qpf
Open

bdarnell wants to merge 2 commits into
tornadoweb:masterfrom
bdarnell:claude/quirky-lamport-0u7qpf

Conversation

@bdarnell

@bdarnell bdarnell commented Oct 7, 2026

Copy link
Copy Markdown
Member

Speeds up _websocket_mask, which runs on every incoming websocket frame and every frame a client sends.

Pure Python (tornado/util.py)

On CPython, the per-byte loop is replaced with a single big-integer XOR: the data and the repeated mask are converted with int.from_bytes, XORed, and converted back with to_bytes, so the per-byte work happens in C. This makes the fallback (used when the extension isn't built) roughly 30x faster for messages of 1 KB or more.

PyPy keeps the existing loop. PyPy's big integers are slow, and its JIT compiles the simple loop well, so the integer version is 2–4x slower on PyPy for anything over about 64 bytes. The implementation is chosen with sys.implementation.name, and both versions are tested on every interpreter.

C extension (tornado/speedups.c)

  • Fix a register-allocation pessimization. data_len had its address passed to PyArg_ParseTuple, and the loop stored through a uint64_t *. uint64_t and Py_ssize_t are the same width (one signed, one unsigned), so C allows them to alias. The compiler therefore stored and reloaded data_len on the stack every 8 bytes. Moving the length into a local whose address is never taken makes large inputs about 3x faster.
  • Remove undefined behavior. Unaligned loads and stores now use memcpy instead of casting char * to uint32_t */uint64_t *. Compilers turn these into the same single instructions, and the code is now valid on strict-alignment platforms. The 8-byte mask is the 4-byte pattern repeated, so the result doesn't depend on byte order. The sizeof(size_t) >= 8 special case is gone; 64-bit arithmetic works on 32-bit platforms too.
  • Lower per-call overhead. METH_VARARGS + PyArg_ParseTuple("s#s#") is replaced with METH_FASTCALL + PyObject_GetBuffer, which saves about 35 ns per call. Both are in the limited API used for the cp311 abi3 wheel. As a side effect, the function now accepts any bytes-like object (bytearray, memoryview) and rejects str; previously it silently UTF-8-encoded str.

I didn't unroll the loop or add explicit SIMD, to keep the code free of compiler- and architecture-specific tuning. Clang vectorizes the simple loop at -O2 anyway; GCC 13 doesn't.

Benchmarks

Per-call time, best of several runs, Linux x86-64 (shared cloud VM, so expect about ±20% noise). CPython 3.13, extension built with the same flags as setup.py (GCC 13, -O2, limited API).

CPython, C extension

size before after
0 B 82 ns 46 ns
7 B 88 ns 58 ns
64 B 92 ns 55 ns
1 KB 188 ns 111 ns
16 KB 1.7 µs 0.68 µs
1 MB 208 µs 76 µs

CPython, pure Python fallback

size before after
0 B 0.33 µs 0.31 µs
7 B 0.93 µs 0.40 µs
64 B 5.0 µs 0.55 µs
1 KB 80 µs 2.7 µs
16 KB 1.32 ms 32 µs
1 MB 82 ms 3.3 ms

PyPy 7.3.23 (3.11), JIT warmed up. No change in behavior; this is why PyPy keeps the loop.

size loop (kept) int XOR C via cpyext
7 B 143 ns 39 ns 394 ns
64 B 231 ns 405 ns 407 ns
1 KB 1.8 µs 6.0 µs 1.5 µs
1 MB 1.66 ms 5.1 ms 0.74 ms

The C extension still isn't worth building on PyPy. Each call through cpyext (PyPy's layer for running CPython extensions) costs about 400 ns, so C only wins above about 1 KB, and then by at most 2x. setup.py already skips it on non-CPython interpreters.

Tests

Added test_lengths, which checks each implementation against a byte-by-byte reference for every length from 0 to 39 (covering all tail cases) and for 255, 256, 1000 and 65537 bytes. Previously the Python fallback was tested only on short inputs and only the version chosen for the running interpreter.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MZ3Gf1f9LuS6xEbDwdPxNB

claude added 2 commits October 7, 2026 18:00
Pure Python: XOR the data against the repeated mask as one big integer
instead of looping over each byte, which is roughly 30x faster for
messages of a kilobyte or more.

C extension:
- Copy the data length into a local whose address is never taken.
  Previously data_len was passed to PyArg_ParseTuple and could alias
  the uint64_t stores, so the compiler reloaded and stored it on every
  iteration (about 3x slower on large inputs).
- Use memcpy instead of casting to uint32_t*/uint64_t*, which was
  undefined behavior for unaligned or differently-typed data.
- Use METH_FASTCALL and the buffer protocol instead of METH_VARARGS
  and PyArg_ParseTuple, which cuts per-call overhead by about 35ns.
  This also stops the function from silently accepting str.

Add a test covering every tail length and some longer inputs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MZ3Gf1f9LuS6xEbDwdPxNB
The big-integer XOR is about 30x faster than the byte loop on CPython,
but on PyPy big integers are slow and the JIT compiles the simple loop
well, so the integer version is 2-4x slower there for anything over
about 64 bytes. Keep both implementations and choose based on the
interpreter, and test both regardless of which one is selected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MZ3Gf1f9LuS6xEbDwdPxNB

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants