Repository navigation
Conversation
Pure Python: XOR the data against the repeated mask as one big integer instead of looping over each byte, which is roughly 30x faster for messages of a kilobyte or more. C extension: - Copy the data length into a local whose address is never taken. Previously data_len was passed to PyArg_ParseTuple and could alias the uint64_t stores, so the compiler reloaded and stored it on every iteration (about 3x slower on large inputs). - Use memcpy instead of casting to uint32_t*/uint64_t*, which was undefined behavior for unaligned or differently-typed data. - Use METH_FASTCALL and the buffer protocol instead of METH_VARARGS and PyArg_ParseTuple, which cuts per-call overhead by about 35ns. This also stops the function from silently accepting str. Add a test covering every tail length and some longer inputs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MZ3Gf1f9LuS6xEbDwdPxNB
The big-integer XOR is about 30x faster than the byte loop on CPython, but on PyPy big integers are slow and the JIT compiles the simple loop well, so the integer version is 2-4x slower there for anything over about 64 bytes. Keep both implementations and choose based on the interpreter, and test both regardless of which one is selected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MZ3Gf1f9LuS6xEbDwdPxNB
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Speeds up
_websocket_mask, which runs on every incoming websocket frame and every frame a client sends.Pure Python (
tornado/util.py)On CPython, the per-byte loop is replaced with a single big-integer XOR: the data and the repeated mask are converted with
int.from_bytes, XORed, and converted back withto_bytes, so the per-byte work happens in C. This makes the fallback (used when the extension isn't built) roughly 30x faster for messages of 1 KB or more.PyPy keeps the existing loop. PyPy's big integers are slow, and its JIT compiles the simple loop well, so the integer version is 2–4x slower on PyPy for anything over about 64 bytes. The implementation is chosen with
sys.implementation.name, and both versions are tested on every interpreter.C extension (
tornado/speedups.c)data_lenhad its address passed toPyArg_ParseTuple, and the loop stored through auint64_t *.uint64_tandPy_ssize_tare the same width (one signed, one unsigned), so C allows them to alias. The compiler therefore stored and reloadeddata_lenon the stack every 8 bytes. Moving the length into a local whose address is never taken makes large inputs about 3x faster.memcpyinstead of castingchar *touint32_t */uint64_t *. Compilers turn these into the same single instructions, and the code is now valid on strict-alignment platforms. The 8-byte mask is the 4-byte pattern repeated, so the result doesn't depend on byte order. Thesizeof(size_t) >= 8special case is gone; 64-bit arithmetic works on 32-bit platforms too.METH_VARARGS+PyArg_ParseTuple("s#s#")is replaced withMETH_FASTCALL+PyObject_GetBuffer, which saves about 35 ns per call. Both are in the limited API used for the cp311 abi3 wheel. As a side effect, the function now accepts any bytes-like object (bytearray,memoryview) and rejectsstr; previously it silently UTF-8-encodedstr.I didn't unroll the loop or add explicit SIMD, to keep the code free of compiler- and architecture-specific tuning. Clang vectorizes the simple loop at
-O2anyway; GCC 13 doesn't.Benchmarks
Per-call time, best of several runs, Linux x86-64 (shared cloud VM, so expect about ±20% noise). CPython 3.13, extension built with the same flags as
setup.py(GCC 13,-O2, limited API).CPython, C extension
CPython, pure Python fallback
PyPy 7.3.23 (3.11), JIT warmed up. No change in behavior; this is why PyPy keeps the loop.
The C extension still isn't worth building on PyPy. Each call through cpyext (PyPy's layer for running CPython extensions) costs about 400 ns, so C only wins above about 1 KB, and then by at most 2x.
setup.pyalready skips it on non-CPython interpreters.Tests
Added
test_lengths, which checks each implementation against a byte-by-byte reference for every length from 0 to 39 (covering all tail cases) and for 255, 256, 1000 and 65537 bytes. Previously the Python fallback was tested only on short inputs and only the version chosen for the running interpreter.🤖 Generated with Claude Code
https://claude.ai/code/session_01MZ3Gf1f9LuS6xEbDwdPxNB