Figures on this page are from the date in the status/header. Performance at the time of writing — re-run on your machine: browser
samples/bench/(live: https://www.o2creative.co.nz/jscolorengine/samples/bench/) or Nodenode bench/mpx_summary.js. Methodology: Bench.md. Canonical tables: BenchResults.md.
Status (2026-08-23). Both caches ship in 1.6. The accuracy-path cache is
src/cache.js(pixelCacheUsed). The in-kernel WASM cache isinterp_*_cachedinside the shipped tetra*.wasm.jsfiles —create()binds it whenpixelCache !== 0(including'auto').kernelInfo().cacheis1when that export ran. Hash-table variants stay in the POC builder only. Two-register double is still a TODO. Historical notes below are labelled.
LittleCMS memoises the last pixel inside cmsDoTransform (unless
cmsFLAGS_NOCACHE): if the incoming pixel is byte-identical to the
previous one, it copies the previous output and skips the
interpolation entirely. jsColorEngine's accuracy path used to be content-neutral in the
same way. WASM 3/4/5/6 kernels now carry a LittleCMS-style last-pixel
compare on the image path (interp_*_cached). JS fallbacks and the
matrix-shaper do not.
That difference is measured and written up in LcmsComparison.md (noise / gradient / solid runs of the native harness). The short version: the cache does nothing on noise, roughly 2–3× on gradients, ~5× on solid fills, and every workflow converges on the same cache-hit ceiling — the cost of compare plus copy with no interpolation at all.
So this is a content-class feature, not a throughput feature.
Kernel3D still leaves the pipeline hint at 'auto' (no accuracy-path
inject on RGB). The image path now binds the in-kernel export on
'auto': a clean photograph is a ~10 % boost; a photograph with
5 % noise added is the worst case (~4 %); solids up to 3.94×.
Node off-vs-auto lives in
BenchResults pixelCache.inKernel.*.
4/5/6 also
keep their accuracy-path inject for transform(). pixelCache: 0
restores the uncached kernel.
Which content classes actually win is not what we assumed going in — the design notes come first here, then what the measurements said, then what shipped.
Still unmeasured — these set the build order, not the outcome.
| Path | Call | Reasoning |
|---|---|---|
| Pipeline / accuracy path | YES — first | Biggest payoff: a hit skips the entire stage walk. See below. |
| 8-bit kernels | YES — test | The typical speed path. Test 1, 2, 16 and 32 slots against real data. |
| int16 | NO | Accuracy-oriented path where the cache suits least, and it doubles the live-register count. Revisit only if the 8-bit data is compelling. |
| SIMD kernels | NO → YES | Reversed once measured: the lanes are channels, not pixels, so there is nothing per-pixel to serialise. 3.07× on flats — see below. |
A hit returns cached output and skips every stage. Two things fall out of that:
Cache at the device boundary, not at the API boundary. Every
dataFormat normalises to a device array (0.0–1.0) in the pipeline's
first step, and converts back in the last. So the cache only ever
needs to handle one input shape — a device array — regardless of
whether the caller passed objects, floats or ints. Check the device
array against the cached one; on a hit, skip straight to the output
block. One implementation covers every format.
This also disposes of the mutability problem: the output converters
all build fresh results (stage_device_to_RGB returns a new literal,
stage_device_to_NCh slices), so the cached device array is never
handed to the caller and cannot be corrupted by them.
Four things to get right, verified against the pipeline builder:
dataFormat: 'device'is the exception. The output converter is gated onconvertOutput && this.dataFormat !== 'device'(Transform.js~3673), so that path has none — a hit would return the cached array itself. Resolved: two variants rather than a runtime test — one that checks and returns a copy ('device'), one that checks and jumps to the output block (everything else).dataFormatis fixed at construction, socreate()picks one of three bound forms — no-cache, cache+converter, cache+device-copy — and the hot path never branches on it.- The output converter isn't reliably last.
insertCustomStage('afterDevice2Output', …)can follow it, andpipelineDebugappends an'END'marker. Jump to the start index of the output block, neverpipeline.length - 1. - Capture indices after
optimisePipeline(), which merges stages and shifts positions. For the same reason, don't hardcode the check at index 0 or 1: the debug'Start'stage only exists whenpipelineDebugis on, andconvertInputmay omit the input converter. Record both indices at build time. - Keep it away from
pipelineDebug. That branch recordspipelineHistoryper stage; jumping would fabricate a history that never ran. Falls out of the branch ordering below at no cost.
The cache check is a normal pipeline stage. The walk reads a step
field instead of incrementing, and the stage sets its own on a hit:
// in the cache stage
stage.step = cacheHit ? stage.stageData.endStep : 1;
// the walk
while (i < len) {
var stage = pipeline[i];
result = stage.funct.call(this, result, stage.stageData, stage);
i += stage.step;
}No new plumbing is needed — funct already receives the stage object
as its third argument (Transform.js ~2629), so a stage can set its
own step today. endStep is the relative jump to the output block,
computed once at create() (post-optimiser, per gotcha 3) and parked
in the stage's stageData.
Give it its own arm rather than changing the shared walk:
if (this.pipelineDebug) { … } // existing, untouched
else if (this.cacheEnabled) { … } // the step-based walk
else { … } // existing i++ walk, untouchedtransform() already branches on pipelineDebug outside the loop and
duplicates the walk, so a third arm is the established idiom, not a
new pattern. Two things fall out of it for free:
- Cache-off pays nothing. The
stepfield replacesi++— a register increment the CPU speculates through — with a load the loop counter depends on, lengthening the loop-carried dependency chain. Probably hidden behind ~50 cycles of per-stage work, but "probably hidden" is exactly the assumption benchmark.md warns about. Confining it to its own arm means the question never arises for anyone who hasn't opted in. - Debug and cache are mutually exclusive by construction, because
pipelineDebugis tested first (gotcha 4, with no explicit guard).
The device/converter split from gotcha 1 collapses into this too: both
are the same stage with a different funct, chosen at create() —
one returns a copy, the other returns the cached device array and
jumps. Same loop, no runtime branch.
The step field is also the general mechanism for runtime-togglable
stages, conditional bypass, or disabling debug without rebuilding the
pipeline, should any of those come up later.
One call site in createPipeline(): after optimisePipeline() but
before the pipeline-validity check (Transform.js ~3686–3691).
After the optimiser so the cache stage can't be folded into a
neighbour; before the verify so its device→device encodings are still
checked — otherwise a malformed injection sails through silently.
if (this.optimise) this.optimisePipeline();
if (this.pixelCache) this.injectCacheStage(); // ← here
// … existing pipeline validity checkEverything the stage needs is fixed at that point, so endStep is
computed once and parked in stageData — never recalculated per
call.
Two injection positions, chosen by input format:
int8/int16→ inject at the front, before the input converter, and key on the raw integers. This is safe rather than merely cheaper: the int→device conversion is deterministic, so identical ints always yield identical device floats — keying on the ints is exactly equivalent to keying on the normalised values, not an approximation of them. Buys an integer key (one exact compare instead of three float compares) and skips the converter on a hit.- All other formats → inject after the converter, letting the
existing object→device stage handle normalisation so the awkward
shapes (
object,objectFloat, arrays) need no per-format cache code at all.
Both positions are resolved at build time, so this is a third build-time axis alongside the variant matrix below — not a runtime branch.
The stage's funct is chosen at build time from a 2×2, so nothing
branches at runtime:
| single entry | keyed (16 / 32) | |
|---|---|---|
dataFormat: 'device' |
check, return copy | check, return copy |
| all other formats | check, jump to output block | check, jump to output block |
Four small hand-written functions — the pipeline path needs no codegen. That idea stays scoped to the kernel matrix, which is the one that actually explodes.
Suggested option shape: a single pixelCache: 0 | 1 | 16 | 32
(0/false = off) rather than separate cacheEnabled + cacheSize,
which admit the invalid enabled: true, size: 0 combination.
The keyed variant needs a quantised index. Unlike the byte
kernels, the pipeline key is a device float array, so there's no
integer to hash. Quantise for bucket selection — (d[0] * 65536) | 0
folded across channels — and keep the exact float compare on the
stored key. Lossy index, exact tag, per the rule above. At ~50 cycles
per stage even a sloppy hash is free here.
Custom stages sitting before the output block are naturally included in the cached value, which is correct — but a side-effecting custom stage (logging, accumulating statistics) would be silently skipped on hits. Either document that or disable the cache when custom stages are present.
Compare the device components directly — don't pack a key. All the
<</& machinery below exists so the byte kernels get one compare per
pixel. Here the stage walk dwarfs any compare, and the device array is
0.0–1.0 doubles that don't pack into an int32 at all. So:
if (d[0] === p[0] && d[1] === p[1] && d[2] === p[2]) { /* hit */ }Exact, no packing code to get wrong, and NaN never compares equal — so a NaN input can never produce a false hit. Fail-safe by construction.
Copy only where there's no output converter. For every format
except dataFormat: 'device', the output block already rebuilds a
fresh result on each hit, so no copy is needed and no stage_*_cacheout
formatter has to be written — the existing converters do the job. The
'device' path has no converter (gotcha 1 above) and hands back the
cached array directly, so that one case must copy.
Where a copy is required, returning the cached array itself is not
merely an aliasing surprise: if the caller mutates it, the cache entry
is corrupted and every subsequent hit returns bad data —
nondeterministic and data-dependent. And skipping the copy buys
little. The accuracy path already allocates ~6 arrays per pixel
walking the stages, so a hit saves ~6 allocations and pays 1 back for
the copy; a no-copy variant reclaims one-sixth of the allocation win
for that bug class. If a read-only caller ever needs it, expose it as
a documented opt-in (unsafeSharedOutput: true) — never the default.
Keep the previous input key and previous packed output in two locals. One compare per pixel.
var prevKey = -1, prevOut = 0; // -1 can never collide: keys are 0..0xFFFFFF
// inside the pixel loop, reusing the r/g/b loads already being done:
var key = (r << 16) | (g << 8) | b;
if (key === prevKey) {
out[o] = prevOut & 255;
out[o+1] = (prevOut >>> 8) & 255;
out[o+2] = (prevOut >>> 16) & 255;
} else {
/* existing cascade, unchanged */
prevKey = key;
prevOut = c0 | (c1 << 8) | (c2 << 16);
}Catches solid fills and flat regions. Two live values, no memory traffic, no allocation. This is the lcms behaviour and the baseline every other option has to beat.
Not a 2-slot hash table. Same shape as (1): two u32 (or two pairs of
i32 locals), two compares instead of one, still no memory. ABAB
dither and a 2-colour logo would hit; ABCABC would not.
Never written, never timed. The paired-export benches jumped from
single-entry to an 8+ slot array in linear memory. That is a
different product (hash, store traffic on every miss). Do not read
those table numbers as a verdict on this variant. Future work: one
inject snippet (_cached_double) next to _cached, same anchors,
same two-export rule — if it does not beat single on logo/dither
without paying more than ~1% on photo, drop it.
Interleaved Int32Array, key at [2n], value at [2n+1], so a hit
touches one cache line:
var idx = (Math.imul(key, 2654435761) >>> (32 - LOG2_SLOTS)) << 1;
// 32 slots → >>> 27, Int32Array(64) 64 slots → >>> 26, Int32Array(128)
// fold the two shifts: (h >>> 25) << 1 ≡ (h >>> 24) & 0xFEIndex on a hash of the whole key, never on a channel. Indexing on
the low byte (blue) collapses to a single slot on any blue-flat
image — sky gradients, single-hue art — giving 100 % conflict misses
and full overhead: strictly worse than (1). 2654435761 is
0x9E3779B1, the nearest prime to 2³²/φ (Knuth multiplicative
hashing, TAOCP Vol. 3 §6.4); the golden ratio is the hardest number to
approximate by a fraction, so clustered keys — exactly what adjacent
pixels are — scatter evenly instead of piling up. Take the high
bits (>>>), never & 255: carries in a multiply only propagate
upward, so the low bits are barely mixed.
Stay small — 32 or 64, not 256. The binding constraint isn't the table's own footprint, it's that the tetrahedral kernel is already streaming a large CLUT through cache. Every line the table occupies is a line not holding CLUT data. 64 slots is 512 bytes / 8 lines; 256 slots is 2 KB / 32 lines. Dither only needs to hold a handful of distinct values locally (a 4×4 ordered dither is ≤ 16), so the bigger table likely buys little hit rate for four times the cache pressure.
- Fuse the key into loads you already do.
(r<<16)|(g<<8)|bis two shifts and two ors against bytes already in registers — no extra memory traffic. This beats aliasing aUint32Arrayover the input, which is impossible for 3-byte RGB anyway: aUint32Arrayview requires a 4-byte-alignedbyteOffset, and pixel n starts at byte 3n. (DataView.getUint32permits unaligned reads but is a bounds-checked call per access — not worth it against three byte loads.) - 32 bits is enough. 3×8 = 24 bits and 4×8 = 32 bits both fit one
int32, so a single
===covers every 8-bit workflow.c3 << 24makes the value negative; that's fine, the comparison is still exact.int16mode is the exception (48 bits, two compares) — see the next section. - No cheap 64-bit in JS.
BigInt64Arrayheap-allocates;Float64Arraybit-compare misbehaves on NaN patterns. InterleavedInt32Arrayis the answer. - Pack the output too, one u32 rather than 3–4 separate cached channel values — it halves the live register count on the hit path.
- Endianness matters only if you store through a
Uint32Arrayview — it's little-endian, so the byte order is the reverse of manual<<packing. Silent-corruption bug if mixed up. - Alpha isn't a complication:
preserveAlphacopies it separately, so the memo only ever covers colour channels.
3×16 = 48 bits doesn't fit an int32, so int16 needs two. The exact
form is cheaper than the 8-bit case, because the channels arrive
as separate Uint16Array elements — you're combining, not extracting:
var k1 = (r << 16) | g; // 2 ops (r<<16 sets the sign bit; harmless, === is exact)
var k2 = b; // 0 ops
if (k1 === prevK1 && k2 === prevK2) { ... }Two ops versus four for int8. The && short-circuits, so a miss —
the common case — usually pays only the first compare. Four-channel
is k1 = (c<<16)|m, k2 = (y<<16)|k, still exact, 4 ops.
No 64-bit needed, and none is available: BigInt64Array
heap-allocates. The f64 trick (r*2³² + g*2¹⁶ + b is exact inside the
53-bit mantissa) was considered and rejected — 4 float ops plus
int→double conversion of the loads loses to two int32 compares.
The real int16 cost is register pressure. Output is 48/64 bits
too, so the memo carries four live values (prevK1, prevK2, prevOut1, prevOut2) against two for int8. With pressure already binding on the
cascade, int16 is the worst candidate for a memo — test it last, if
at all.
Use a uniform 4-word slot, decided by cache-line alignment rather
than size. int16 RGB is 48-bit key + 48-bit result = 96 bits, which
packs into exactly 3 words with no waste ([key R,G][key B | result B][result R,G]), and int16 CMYK is 64+64 = 128 bits = 4 words with no
masking at all. Tempting to use 3 for RGB — but a 64-byte line is 16
words, so 4-word slots divide a line exactly (4 per line, never
straddling), while 3-word slots land at word offsets 0, 3, 6, 9, 12,
15… and cross a line boundary on roughly one access in five. The
saving is 256 bytes at 64 slots, in a table that is L1-resident
either way; the straddle is not worth it. Masking cost is the minor
consideration (~2 ops), not the deciding one.
int8 doesn't raise the question: 24+24 and 32+32 bits both fit 2 words, which is the interleaved layout above and is equally self-aligning at 8 slots per line.
Tempting, because it's true that the kernel doesn't actually compute
16 bits of precision: the int16 POC keeps Q0.8 frac, so only 8 of
the 16 fractional bits weight the interpolation (bench/int16_poc/ RESULTS.md, "Accuracy ceiling"). So a 16-bit-exact key is arguably
over-precise relative to what the kernel produces, and masking down to
10+10+10 (or the display-style 10+11+10) fits one int32 again.
Rejected on four counts:
- It's slower.
((r>>>6)<<20) | ((g>>>6)<<10) | (b>>>6)is 7 ops against the exact version's 2. Lossiness doesn't buy op count — it buys hit rate. - It buys that hit rate on exactly the wrong content. Lossy keys only help on smooth gradients, which is the one thing 16-bit mode exists to serve. The feature would degrade the workload that asked for it.
- The error isn't a clean quantisation. A bucket returns whatever pixel landed in it first, so output becomes scan-order dependent: the same image cropped differently converts differently. That is non-determinism, not precision loss.
- It's unverifiable. Content- and order-dependent output can't be checked against the lcms oracle, which breaks the accuracy story and the v1.6 QC plan.
(The 10+11+10 split gives green the extra bit for luminance sensitivity — a display convention. It doesn't transfer here: the error propagates non-linearly through the CLUT rather than landing on the eye.)
Two places lossiness is legitimate:
- As the hash index for the 32/64-slot table. A bucket selector is
allowed to be lossy; only the key stored in the slot must be
compared exactly.
idx = hash(high bytes)is fine and cheap. - As an explicit input pre-quantise. If the speed is wanted, round the input to N bits up front and then run the exact memo. Same hit-rate win, fully deterministic, oracle-verifiable, and the caller opted in knowingly — rather than a hidden lossy compare inside the kernel.
| Cost | Applies to | Notes |
|---|---|---|
| ~6 ALU ops per pixel on a miss | all | Against ~40–60 ops for the 3D int cascade — estimated 8–12 % tax on photographic content; proportionally less on 4D. |
| Register pressure | (1), (2) | Cached values stay live across the cascade, where JitInspection.md already found pressure binding. V8 will spill. This is the main reason to consider the run-scan restructure below. |
| Branch misprediction | all | Solid and noise both predict perfectly (always / never taken). Dither is the bad case — alternating outcomes, ~15–20 cycles a pop. Could exceed the op-count estimate. |
| Store traffic on every miss | (3) only | Photographic content misses on every pixel, so every pixel writes two slots. (1) and (2) update registers — free. This is what most likely sinks the table on photos. |
| CLUT cache-line eviction | (3) only | See sizing note above. |
Restructure that removes the register pressure: identical semantics to (1) — both only ever catch consecutive identical pixels, both produce bit-identical output — but the cache state never sits live across the cascade:
while (i < n) {
var key = pack(in, p);
var j = i + 1, q = p + 3;
while (j < n && pack(in, q) === key) { j++; q += 3; }
interpolateOnce(key, tmp); // hot kernel, zero cache state
fillRun(out, i, j, tmp); // tight store loop
i = j; p = q;
}On noise, runs are length 1 and the cost matches (1). On flat content
fillRun is pure stores and should run far past the lcms compare-and-
copy ceiling, because lcms still pays a per-pixel compare inside a run
and this doesn't.
Everything in this section is 3- and 4-channel. Above 4 input channels the LUT is declined and the arithmetic changes — see N-channel input is a different regime.
Working hypothesis, and the thing to test first:
A cache's payoff scales with the cost of the work it skips. The no-LUT accuracy path runs ~6–11 MPx/s, walking the full stage pipeline with per-pixel allocation — a hit there skips two orders of magnitude more work than a hit in the int kernel, while the ~6-op key cost stays constant. Register pressure is a non-issue on that path (it is already allocating arrays per pixel), and the branch is noise against the stage walk.
The tuned LUT kernels are the opposite: 49–73 MPx/s, register pressure already binding, and ~6 ops is a much larger relative tax. They may simply be fast enough that the cache is a pure cost.
This inverts the usual instinct to optimise the hottest loop, and it is cheap to check — so check it before building anything for the kernels.
Never in the SIMD kernels — reversed below. The exclusion held only while the lanes were assumed to be pixels.
Three statements above rule the cache out of the SIMD kernels on the grounds
that "a scalar check serialises what the f32x4 path vectorises". That is a
correct statement about a pixel-parallel kernel, and tetra3d_simd is not one.
The lanes are the four u16 channels at a CLUT corner, picked up by a single
v128.load64_zero + i32x4.extend_low_i16x8_u. One loop iteration is ONE
PIXEL. Pixel-parallel was in fact tried first and lost — 0.89×, because each
lane needed its own LUT gather — which is exactly why the axis was flipped, and
WasmKernels.md
records it. So there is no pixel-level vectorisation for a per-pixel check to
serialise. The cache drops into the SIMD kernel as easily as into the scalar
one, on the path that actually ships.
The pull is easy to keep feeling: "SIMD" primes you for four pixels at a time, and here it means four channels of one pixel.
Measured, paired exports against the shipped tetra3d_simd, all outputs
byte-identical:
| content | cached ÷ shipped |
|---|---|
| solid | 3.07× |
| logo, 5% mark on white | 2.40× |
| logo, 30% mark on white | 2.57× |
| ILLUSTRAT | 1.04× |
| photographs | 0.93–0.96× |
| noise | 0.99× |
Scalar hits keep a $prevOutPtr to the last colour write — copying
from outputPos - cMax would pick up the previous pixel's alpha
byte. SIMD already holds colour in $prevOut, so it did not need
that.
Alpha is not in the key, and must not be. tetra3d reads three bytes,
advances inputPos by three, and handles alpha in a tail; the cache brackets
only the colour work, so the tail runs on a hit as well as a miss. A solid RGB
under a per-pixel alpha gradient therefore hits every pixel — measured at
2.80x, byte-identical, alpha preserved exactly. Keying on the whole RGBA
register would have scored that image at 0%.
And it is the single-entry cache that wins, not a hash table. Not "have I seen this colour before" — "did the same bits arrive as last time, so the same bits can leave". The previous key is an i32 local and the previous output is the v128 the kernel was about to store anyway, so a hit is one compare and one store, with no memory touched and no parameters added. It never interprets the pixel, which is why one insertion covers int8, int16, RGB and RGBA, and why the scalar and SIMD versions are the same twenty lines.
A 4096-entry hash table was measured beside it and is worse everywhere except photographs, where it needs 32 KB to reach 0.95–2.5× — and photographs are the content this whole feature is not for.
Two design conclusions that only measurement produced:
- Not a runtime mode. Behind a
$cacheModeparameter the UNCACHED path measured 15–22% slower, and got worse when a third mode was added — the cost is the code behind the guard, not the guard. A single mode compare was worth ~10% on its own: swapping which mode was tested first swapped which one won. - Paired exports instead. One module, two functions,
interp_tetra3d_simdandinterp_tetra3d_simd_cached, the uncached one copied in verbatim. It measures 0.985–1.008× against the shipped binary — a tie, as it must be, since there is no cache code in it to pay for. Enabling the cache is swapping a function reference; the signature is identical.
scripts/compile_kernel_wat.js injects the single-entry twin
(interp_* + interp_*_cached) into every tetra *.wasm.js.
create() swaps the function reference when pixelCache !== 0.
POC benches and table variants stay in bench/pixel_cache_wasm/.
Two-register double: § TODO.
The combinatorial problem: cache off / single / 2-entry / 32 / 64,
× dimension, × output channels, × lutMode. Hand-writing that matrix is
untenable, and a runtime if (cacheEnabled) inside the loop taxes the
no-cache path — so today it would have to be all-or-none.
Codegen resolves it: emit the loop with the cache config baked in as literals, so the no-cache variant contains no cache code at all and the table size is a constant, not a load. Points in its favour here:
- The emitter infrastructure already exists —
emit_js_*/attachStore_js_*insrc/stages.js, andcompile()/getSource()/toModule()in CompiledPipeline.md. There is already anew Functionrunner experiment queued in benchmark_todo.md. - A generated kernel registers as a normal kernel descriptor
(
lutMode: 'int-memo', etc.), so it can be A/B'd against'int'with zero risk to the default path or any published number. - Each generated variant is a fresh function object with its own type feedback — monomorphic per config, no polymorphic dispatch. That is a genuine advantage, not just a packaging convenience.
Costs to keep in view: compile plus tier-up warmup per generated
variant (cache generated functions by config key so repeated identical
configs share one object), and new Function is blocked under a
strict CSP without unsafe-eval — a browser library must fall back
to the hand-written kernels, so codegen can never be the only path.
For the cache that fallback is trivially correct: no codegen means no
cache, which is the default anyway.
bench/pixel_cache_kernel_poc/ builds its cached kernel exactly this
way (toString() → insert → new Function) and produces byte-identical
output, so the mechanism is no longer hypothetical.
More importantly it resolves a problem hand-writing cannot: both the injection point and the key method vary by data type, and neither can be a runtime test without taxing the hot loop.
| input | key | where to inject |
|---|---|---|
| u8, 3 ch | (r<<16)|(g<<8)|b — 24 bits, one int32 |
straight after the input reads |
| u8, 4 ch | (k<<24)|(c<<16)|(m<<8)|y — exactly 32 bits, one int32 |
straight after the input reads |
| u16, 3 ch | 48 bits → two int32, &&-chained compare |
after the input reads |
| u16, 4 ch | 64 bits → two int32 | after the input reads |
| float | no packing — compare components, or hash raw bits | after normalisation, not before |
Baking that in as literals means the emitted loop carries one key expression and one compare with no branching, and the uncached variant contains no cache code at all. The earlier note that scoping to 4D/ND made codegen unnecessary was too quick: two dimensions × three data types × channel counts × slot sizes is exactly the matrix codegen exists for.
Inject into the hand-written kernels; do not generate them from scratch. The kernels are the product of the tuning recorded in JitInspection.md and the PERFORMANCE LESSONS block, and that hand-work is the asset. Transforming their source at build time keeps them the single source of truth and guarantees the cascade stays identical — which is precisely why the POC's output matched byte for byte on the first run.
Hit rate, and it needs no kernel work at all — the pipeline cache is the instrument. Build that first (it's shipping anyway), give the test build a hit counter, and run a real corpus through it. The measurement effort isn't throwaway, unlike a standalone counting script.
Three things make it predictive of the kernel decision:
- Feed it images in scanline order, not swatches. Hit rate is a
property of the data; the pipeline cache only predicts the kernel's
rate if it sees the same pixel sequence the kernel would. Use
transformArraywithbuildLut: false. - Speed of the instrument is irrelevant. Hit rate converges after a few hundred thousand pixels, so 6–11 MPx/s is ample — no need for full-resolution runs.
- Only the hit rate transfers — never the timings. The cost sides are completely different: register pressure and branch misprediction dominate in the kernels and are near-free on the pipeline path. Accuracy-path speed numbers are not a kernel verdict.
Count per image: single-entry hits, 2-entry hits, and 16/32-slot table hits. That alone decides which shape (if any) is worth building.
Corpus must include:
- photographs (the case that pays the tax)
- UI screenshots and flat vector art (the case that pays out)
- dithered / halftone content — otherwise the bench looks better than reality, since that is the branch-mispredict case
- the noise / gradient / solid synthetics, mirroring
BENCH_INPUTin the native harness (bench/lcms_c/) so results line up with the lcms measurements
Then, only if the hit rate justifies it: the miss-path tax on the accuracy path and on the int kernel, measured separately.
One prior observation worth re-checking: our 4D kernel measured ~20 % slower on uniform content than on noise, so the flat-art win starts from a slightly worse baseline than the noise numbers suggest.
node bench/pixel_cache/cache_bench.js — sRGB → AdobeRGB,
buildLut: false. Hit rate is the transferable figure; the MPx/s
columns describe this path only.
| content | cache off | slots=1 | slots=32 |
|---|---|---|---|
| noise | 7.75 | 0.0 % · 0.82x | 0.0 % · 0.82x |
| gradient | 9.01 | 75.0 % · 1.85x | 75.0 % · 1.78x |
| checkerboard | 9.06 | 0.1 % · 0.76x | 100 % · 3.25x |
| palette8 | 9.10 | 12.5 % · 0.86x | 87.5 % · 1.94x |
| solid | 8.83 | 100 % · 3.39x | 100 % · 3.25x |
| near-miss (1 LSB apart) | — | 0 % | 83 % |
Unsplash photographs and one Library of Congress poster, 7.6–19.4 MP, every pixel converted:
| image | pixels | cache off | slots=1 | slots=32 |
|---|---|---|---|---|
| beach | 11.9 M | 4.92 | 1.0 % · 0.82x | 3.2 % · 0.84x |
| sunflower | 7.6 M | 8.38 | 10.8 % · 0.85x | 19.4 % · 0.87x |
| strawberries | 10.8 M | 8.74 | 25.4 % · 0.93x | 36.9 % · 0.99x |
| photo of text page | 18.7 M | 4.46 | 13.0 % · 0.94x | 41.5 % · 1.05x |
| poster illustration | 19.4 M | 4.60 | 31.3 % · 1.01x | 67.4 % · 1.22x |
Conclusion: the original hypothesis was right. Photographs run 3–41 % and land between break-even and a 16 % loss. The one clear win is the flat-colour illustration at 67 % and 1.22x. The cache is a content-class feature that pays on graphic and synthetic content and costs on photographic content, exactly as the design notes argued.
Recorded because both were convincing while they lasted.
1. The bundled sample images are not photographs. A first pass over
samples/images/ (face, fruit, skin) showed 59–83 % hit rates
and 1.2–1.9x, and this document briefly claimed the photograph
assumption had been "overturned" — that a keyed table catches
recurrence rather than adjacency and therefore wins on photos too.
Those three are AI-generated/adjusted images with large flat
backgrounds and unnaturally smooth gradients. On real camera output the
effect largely disappears. The recurrence-vs-adjacency mechanism is
real (the near-miss row above shows 0 % → 83 %); what was wrong was
the claim that natural photographs supply enough of it.
2. Capping pixels crops rather than samples. --pixels takes the
first n pixels, which on a 19 MP frame is the top 1–3 % — usually
sky or background, and nothing like the whole image. The beach photo
reads 27.8 % resized-to-small, 7.5 % as a 250k top crop, and
3.2 % over the full frame. Striding would sample evenly but destroy
adjacency, which is precisely what the slots=1 column measures, so
the bench now warns when it crops and the answer is to raise
--pixels, not to stride.
Still not a corpus. Five images. Missing: screenshots and UI captures, halftone/dithered content, and scanned or print-origin material — the classes most likely to favour the cache. And note the uncached column is itself content-sensitive (4.46–8.74 MPx/s), so the pipeline was never truly content-neutral either.
Noise never hits, so its throughput is the tax on its own:
| MPx/s | vs off | key memory | |
|---|---|---|---|
| cache off | 7.91 | — | — |
| slots=1 | 6.46 | −18.4 % | — |
| slots=32 | 6.17 | −21.9 % | 1 KB |
| slots=256 | 5.93 | −25.1 % | 6 KB |
| slots=1024 | 5.98 | −24.3 % | 24 KB |
| slots=4096 | 6.05 | −23.4 % | 96 KB |
The cost is the check, not the table. Going from 32 to 4096 slots — a 96× larger working set — changes nothing measurable; the numbers past 256 sit inside run-to-run noise. What costs is doing a lookup at all: the extra stage call, the hash, and the compare. Single-entry is cheaper (−18 %) only because it skips the hash and indexing.
The L1-pressure worry in the design notes above does not apply here,
and the reason is structural rather than a matter of degree: this is a
full JS pipeline that allocates several arrays per pixel and dispatches
every stage through .call(). Memory traffic and allocation already
dominate, so a few KB of table is invisible. A hot kernel is the
opposite regime — tight integer loops streaming a large CLUT with
nothing to hide behind — and there the same table genuinely would
compete. The argument was imported from the wrong path, not merely
overstated.
~18 % looks steep for "a hash and a lookup", so it was decomposed by adding plain pass-through stages to the same pipeline:
| MPx/s | vs baseline | |
|---|---|---|
| baseline, 7 stages | 7.80 | — |
| + 1 no-op stage | 7.56 | −3.1 % |
| + 2 no-op stages | 7.29 | −6.5 % |
| cache slots=1 (no hash at all) | 6.44 | −17.4 % |
| cache slots=32 (hashed) | 6.29 | −19.4 % |
Roughly: 6.5 points is bare stage dispatch, ~8 points is
bookkeeping (counters, key writes, stage.step writes, copying the
value), and the hash is about 3 points — the gap between slots=1,
which does no hashing whatsoever, and slots=32.
So the hash is the cheapest part, and the dominant cost is structural:
in this architecture every stage is a dynamic .call() passing arrays
around, so adding two stages costs 6.5 % before they do anything.
That is the price of the plugin shape — the cache can be declined,
switched off at zero cost, and deleted without trace, precisely because
it is ordinary stages rather than something welded into the walk.
One fix came out of this: the store stage originally did
colour.slice(), allocating on every miss — the common case for
photographic content. Reusing the slot's array took slots=1 from
−17.4 % to −14.8 % and slots=32 from −19.4 % to −17.5 %. Worth having,
but it confirms there is no single large win hiding here; the cost is
spread thin.
Further reduction would mean a build-time variant with the counters compiled out, or folding the check into the walk instead of using stages — trading the properties above for a couple of points.
The hash itself was timed in isolation at 1.35 ns/pixel for three channels. Restructuring it as three independent multiplies combined at the end — on the theory that the chained form is latency-bound — came out slower (1.50 ns). It is not a dependent-chain problem, and there is nothing to win there. Recorded so nobody tries it twice.
This changes the earlier conclusion, which was too pessimistic. An
unrolled kernel loop has no stage dispatch at all, so the 6.5 points
that dominate here simply do not exist, and most of the bookkeeping
(the stage.step writes, the separate store stage, the counters) goes
with it. The check reduces to: build the key from bytes already loaded
for interpolation, hash, compare, branch — roughly 4–5 ns.
Against that, the per-pixel budget shrinks: the int 3D kernel runs at ~73 MPx/s (13.7 ns/pixel) and the 4D at ~48 MPx/s (20.8 ns), versus ~125 ns here. A rough model — estimates from decomposition, not measurements:
| miss cost | hit cost | break-even | ceiling | |
|---|---|---|---|---|
| accuracy path (measured) | 1.21× | 0.31× | ~38 % | 3.2× |
| 3D int kernel (modelled) | ~1.33× | ~0.52× | ~41 % | ~1.9× |
| 4D int kernel (modelled) | ~1.22× | ~0.34× | ~25 % | ~2.9× |
So the kernel port is not obviously a loser: break-even lands in much the same place, and the 4D/CMYK kernel is the most attractive target of all — the heavier the per-pixel work, the better the cache looks, exactly as the pipeline-weight measurements show on the accuracy path.
Two caveats before anyone builds it. These are models, not measurements — the honest next step is a POC on one kernel rather than more arithmetic. And register pressure is unmodelled: the cache state must stay live across the interpolation cascade, where JitInspection.md already found pressure binding, which is exactly what the run-scan restructure earlier in this document exists to address.
Built and measured: bench/pixel_cache_kernel_poc/, a DeviceLink
CMYK→CMYK against tetrahedralInterp4DArray_4Ch_intLut_loop, with the
cached variant produced by source-transforming the real kernel so the
cascade is identical to production.
| content | plain | cached (64 slots) | change | hit rate |
|---|---|---|---|---|
| cmyk noise | 26.9 | 24.6 | −8 % | 0 % |
| photo → CMYK | 31.8 | 46.6 | +47 % | 57 % |
| photo → CMYK | 27.8 | 41.5 | +49 % | 53 % |
| poster → CMYK | 33.0 | 59.6 | +81 % | 73 % |
| cmyk gradient | 32.7 | 72.2 | +121 % | 75 % |
| cmyk solid | 32.3 | 151.4 | +368 % | 100 % |
Byte-identical output in every case. Break-even is ~10 %, against ~38–40 % on the accuracy path — so the modelled ~25 % was pessimistic, and the kernel port it made look questionable is comfortably worth doing.
Why it is so much better:
- No stage dispatch — the largest single component of the accuracy-path tax does not exist in an unrolled loop. Miss tax 8 %.
- The key is one int32. 4 × u8 packs exactly, so the check is a
single
===and the hash a singleimul, against three float compares and three chained imuls. - A hit skips proportionally more — the tetrahedral cascade is the bulk of the per-pixel cost.
The most important finding is about the data, not the code. The same photographs that hit 3–41 % as RGB hit 43–70 % once separated to CMYK, because the RGB→CMYK LUT quantises many source colours onto the same output. So CMYK wins twice over: heavier per-pixel work and inherently more repetitive data. That is the strongest argument yet for scoping the cache to CMYK/4D.
Table size matters here, unlike on the accuracy path, because the kernel streams a large CLUT and competes for L1: 16–256 slots all cost −7 to −8 %, but 1024 (12 KB) drops to −17 %. Use 64–256. The L1-pressure reasoning from the design notes holds here; what it does not carry to is the allocation-heavy accuracy path, where a few KB of table is invisible.
Caveats: 8-bit only (the 32-bit key does not generalise to u16 or float), no alpha exercised, one kernel on one machine. A POC, not a feature.
Worth reading the reference implementation once the design was settled,
to see where two independent attempts agreed. CachedXFORM in
cmsxform.c:
typedef struct {
cmsUInt16Number CacheIn [cmsMAXCHANNELS]; // 16 ch x u16 = 32 bytes
cmsUInt16Number CacheOut[cmsMAXCHANNELS]; // 32 bytes -> 64 total
} _cmsCACHE;
accum = p->FromInput(p, wIn, accum, ...); // unpack to u16
if (memcmp(wIn, Cache.CacheIn, sizeof(Cache.CacheIn)) == 0)
memcpy(wOut, Cache.CacheOut, sizeof(Cache.CacheOut));
else { p->Lut->Eval16Fn(wIn, wOut, p->Lut->Data);
memcpy(Cache.CacheIn, wIn, sizeof(Cache.CacheIn));
memcpy(Cache.CacheOut, wOut, sizeof(Cache.CacheOut)); }
output = p->ToOutput(p, wOut, output, ...); // packWhere we agreed, independently:
- The cache sits between unpack and pack — on the unpacked pixel, with format conversion still running every pixel. That is exactly the device-boundary position argued for above.
- Variants are selected at transform creation, not tested in the loop:
lcms swaps between
PrecalculatedXFORM,CachedXFORMand the two gamut-check equivalents via a function pointer. - Float transforms get no cache at all (
dwFlags |= cmsFLAGS_NOCACHE).
Where we differ:
- lcms caches one entry only. No table, no hash. Our measurements say that is the weak form — single-entry managed 1–31 % on photographs where a keyed table reached 43–70 % in CMYK — so the table is a genuine step past the reference, not a re-implementation of it.
- Fixed-size
memcmpover all 16 channels, which is whywInis zeroed first: unused channels are zero on both sides, so one constant-size compare (two SIMD ops) serves any channel count with no loop and no branch. We specialise per channel count instead. - The cache is stack-local per call (
memcpy(&Cache, &p->Cache, …), never written back), making it thread-safe by construction. Ours persists on the Transform — better across many small calls, irrelevant for one whole-image call.
What we took: seeding. lcms initialises its entry at transform
creation by evaluating the all-zero pixel, so the cache always holds a
real (key, value) pair and has no empty state at all. Adopted here,
and it earns more than tidiness:
- The
slotValidarray is gone — one load removed from every lookup. - The single-entry
hasPreviousflag is gone. Measured:slots=1improved on every content class (skin.png 1.03× → 1.25×, fruit 1.20× → 1.24×, noise 0.82× → 0.87×). - In the kernel it removes a real problem. With a packed 4×8-bit
CMYK key there is no impossible int32 sentinel — CMYK(255,255,255,255)
is exactly
-1— so the first POC had to keep keys in aFloat64Arrayinitialised toNaN. Seeded, keys go back toInt32Array: a quarter of the memory (256 slots: 1 KB vs 4 KB) and an integer compare instead of a double one.
Filling every slot of a table with the same seed pair is safe, not just the slot it hashes to: a given key only ever probes one slot, so duplicates elsewhere are unreachable by the seed key, and any other key landing there compares unequal and misses. The copies are simply overwritten as real entries arrive.
The cost is one extra key write per miss. With no validity flag a half-written entry is indistinguishable from a good one, so the key can no longer be written by the check and the value by the store — a stage throwing between the two would leave a new key beside a stale value and silently corrupt every later hit on it. Both halves are now written together in the store stage, which costs about 1 point on the pure-miss path and is worth it.
setPixelCacheProfiling(true) answers "is this worth enabling for my
content?" without the caller needing to know anything about the
internals.
The implementation is a single number. A miss normally sets
stage.step = 1 and walks on into the maths; in profiling mode it sets
endStep - 1, jumping straight to the STORE stage. No NOP stage, no
splice, no second pipeline — the pipeline is untouched and the mode is
a runtime toggle.
Hit accounting stays bit-identical, and the reason is structural: the key depends only on the stages before the check, and those still run. Only the value-producing stages are skipped. Measured on palette content, normal and profiling report the same hits and lookups to the integer.
Speed-up is inversely proportional to hit rate, which is the useful way round:
| content | hit rate | normal | profiling | speed-up |
|---|---|---|---|---|
| noise | 0 % | 6.2 | 18.8 | 3.0× |
| gradient | 75 % | 14.2 | 24.8 | 1.8× |
| solid | 100 % | 26.9 | 27.5 | 1.0× |
Profiling is fastest exactly when the content is worst for the cache — i.e. when the answer is "don't bother" and you want it quickly. When it is slow, the normal run was already fast and you have your answer anyway.
The output is meaningless while it is on. The cached values are whatever was in flight, not converted colour, so the cache is flushed on every mode change in both directions. Without that, profiling and then transforming for real would silently return garbage from every hit — which is why the toggle owns the flush rather than leaving it to the caller.
Targeted RGB/CMYK caching on the common 8-bit path — not a generic cache. JavaScript performance is won by narrowing, not by generality, so the cache goes only where it demonstrably pays and the affected surface stays as small as possible:
| in scope | file |
|---|---|
tetrahedralInterp3DArray_3Ch_intLut_loop |
src/kernels/3d/kernel3D_loops.js:269 |
tetrahedralInterp3DArray_4Ch_intLut_loop |
src/kernels/3d/kernel3D_loops.js:665 |
tetrahedralInterp4DArray_3Ch_intLut_loop |
src/kernels/4d/kernel4D_loops.js:900 |
tetrahedralInterp4DArray_4Ch_intLut_loop |
src/kernels/4d/kernel4D_loops.js:1149 |
Four loops. Everything else is explicitly out: 1D and 2D (their input space is enumerable — precompute instead, see below), ND (float, proof-oriented, and the key does not pack), all float variants, every WASM and SIMD variant (a scalar check serialises what f32x4 vectorises), and the matrix-shaper path (headed for ~4 ns/pixel, where a 4–5 ns check would double the cost).
The SIMD exclusion did not survive measurement. It rests on the lanes being pixels; in
tetra3d_simdthey are channels, so there is nothing per-pixel to serialise — 3.07× on flats. Worked through in SIMD here is parallel interpolation, not parallel pixels. The matrix-shaper exclusion does stand, now measured at 3.0 ns/pixel.
These four are chosen because they are the real-world common paths —
8-bit RGB and CMYK image conversion — and because 8 bits × 3 or 4
channels packs into a single int32, which is what makes the check one
=== and one imul. That property is the whole reason this is cheap;
it does not survive to u16 or float, and neither should the feature.
Still to measure: the 3D case. Only 4D has a POC
(break-even ~10 %, +47 % to +169 % on real content). 3D was modelled at
~41 % break-even and never measured, and RGB content hit rates are more
variable than CMYK — beach photo 3 %, text page 41 % at 32 slots. The
same POC should be run on tetrahedralInterp3DArray_4Ch_intLut_loop
(RGB → CMYK, a heavily used soft-proof and print-prep path with
4-channel output) before 3D ships alongside 4D.
Not all of them, and the rule is sharper than "the heavy ones":
A cache is only interesting where the input space is too large to precompute and the per-pixel work is high.
At 8 bits per channel the whole input space is enumerable for small dimensions, and a complete table beats any cache:
| input | distinct inputs | verdict |
|---|---|---|
| 1D (Gray) | 256 | Never cache. Precompute all 256 outputs — that is just a LUT, and cheaper than one hash. |
| 2D (Duotone) | 65,536 | Never cache. A complete table is 64 K entries; precompute it. |
| 3D (RGB) | 16.7 M | Marginal. Too big to enumerate (this is the 2²⁴ joke above), but per-pixel work is only moderate — modelled break-even ~41 %. |
| 4D (CMYK) | 4.3 G | The target. Cannot be enumerated, heaviest interpolation, modelled break-even ~25 %. |
| ND (5–15 ch) | astronomical | Also attractive — KernelND runs the per-pixel interpolator with an allocation per pixel, so it is the slowest path in the engine. |
Output channel count pushes the same way: 4 output channels is more interpolation per pixel than 3, so CMYK and N-channel destinations improve the ratio too. A DeviceLink CMYK→CMYK is the ideal case — 4D in, 4 out, and print workflows are exactly where flat graphic content lives.
RGB matrix-shaper needs no cache, now or later. That path is headed for the fused WASM kernel at 250–257 MPx/s (MatrixShaperKernel.md) — about 4 ns per pixel. A 4–5 ns check would more than double the cost of the thing it was meant to accelerate.
Which is the general warning: the better a kernel gets, the worse the cache looks on it. The two efforts pull against each other, so the cache belongs only on paths that are expensive for irreducible reasons — high-dimensional interpolation — not on paths that are merely unoptimised yet.
Scoping to 4D/ND does not remove the case for codegen, as first thought. Even two dimensions still multiply out across data type (u8 / u16 / float), channel count and slot size — and crucially the injection point and key method differ per data type, which cannot be a runtime test without taxing the loop. See "Validated by the POC" under Codegen above.
| image | 1 | 16 | 32 | 64 | 128 | 256 | 1024 |
|---|---|---|---|---|---|---|---|
| text page | 13.0 % | 35.2 % | 41.5 % | 47.7 % | 56.0 % | 63.7 % | 79.6 % |
| strawberries | 25.4 % | 35.3 % | 36.9 % | 38.7 % | 40.4 % | 41.8 % | 48.7 % |
| poster | 31.3 % | 62.3 % | 67.4 % | 70.5 % | 73.5 % | 76.9 % | 83.6 % |
| sunflower | 10.8 % | 17.5 % | 19.4 % | 21.3 % | 22.9 % | 24.6 % | 33.8 % |
| beach | 1.0 % | 2.4 % | 3.2 % | 4.4 % | 6.0 % | 8.0 % | 13.0 % |
Hit rate never plateaus. It is still climbing at 1024 slots on every image — these are 8–19 MP frames, so even 1024 entries is a tiny window over the colours present. Since the cost is flat, a bigger table is strictly better, bounded only by memory you care about.
Counting distinct colours per image explains why the numbers above are so modest — and it is not because photographs lack repetition:
| image | pixels | unique colours | perfect-cache ceiling | actual @32 | @1024 |
|---|---|---|---|---|---|
| text page | 18.7 M | 0.02 M | 99.9 % | 41.5 % | 79.6 % |
| poster | 19.4 M | 0.18 M | 99.1 % | 67.4 % | 83.6 % |
| strawberries | 10.8 M | 0.47 M | 95.6 % | 36.9 % | 48.7 % |
| sunflower | 7.6 M | 0.31 M | 96.0 % | 19.4 % | 33.8 % |
| beach | 11.9 M | 0.88 M | 92.6 % | 3.2 % | 13.0 % |
Every image could hit 92–99.9 %. Even the beach photo repeats itself constantly — 11.9 M pixels drawn from 0.88 M colours. What the cache actually delivers is 3–67 %, so the losses are almost entirely conflict misses in a small direct-mapped table, not a shortage of repetition in the data. Associativity or a much larger table would recover far more than tuning anything else here.
Verified independently: a standalone direct-mapped simulation, sharing no code with the engine, reproduces the engine's hit rate to the digit on every image at both 32 and 1024 slots.
Which makes the limit case obvious. Hit rate never plateauing does suggest an easy fix: just use 16,777,216 slots and cover every possible 8-bit RGB colour. 100 % hits after first touch, problem solved. 🎉
That is, of course, a lazily-populated 256³ CLUT — which is
buildLut: true, built eagerly, packed as u16, and with a decade of
kernel work behind it. Congratulations, you have reinvented the LUT,
slower and one pixel at a time.
Which is the real framing for this whole feature: the pixel cache is
a partial, lazy LUT for the path where a full one is not wanted. Grow
it far enough and it becomes the thing the engine already does better.
That also bounds how much effort it deserves — anyone needing high hit
rates on bulk data should be using buildLut, not a bigger cache.
All the figures above use sRGB → AdobeRGB, which at 7 stages (two
gammas and a matrix) is the cheapest pipeline in the engine — so the
fixed cost of a cache check is at its most visible. Heavier pipelines
dilute the tax and amplify the win:
| pipeline | noise (pure tax) | poster @32 | poster @1024 |
|---|---|---|---|
| sRGB → AdobeRGB (7 stages) | −19 % | +112 % (85 %) | +191 % (99 %) |
| sRGB → GRACoL (10 stages, 3D CLUT) | −15 % | +156 % (85 %) | +269 % (99 %) |
So on a CMYK destination with graphic content the cache is worth 3.7×, against 2.9× for the same content on the cheap RGB pipeline. Break-even moves down accordingly. Anyone measuring this on their own content should measure it on their pipeline — an RGB→RGB result is the pessimistic end of the range.
(A 4-channel input also makes the check itself dearer — four values
to hash and compare instead of three — which is why GRACoL → sRGB
shows −19 % rather than following the dilution trend.)
A hit runs at roughly 0.3× the cost of an uncached pixel and a miss at roughly 1.25×, which puts break-even near 38–40 % hit rate. The measurements land exactly there: strawberries at 36.9 % gives 0.99×, text page at 41.5 % gives 1.05×.
So the decision is on or off, not what size: pick the largest table you are willing to pay memory for, then ask whether the content clears ~40 %. Single-entry has a lower break-even (~24 %) but a ceiling so low that it only wins on near-solid content.
Verified, not assumed. bench/pixel_cache/verify_cache.js compares
cached against uncached output byte for byte — every content type
above, four transform shapes (3- and 4-channel output, int8 and int16),
all cache modes: 108 whole-image comparisons, every byte identical,
FNV-1a hashes matching. A reduced version runs in the test suite. This
matters because colour-level unit tests cannot generate the evictions,
collisions and hit/miss interleavings that only appear at image scale.
Worst case is worse than estimated. Pure noise costs 0.82x — about
a 20 % tax, against the 8–12 % predicted from op counts. And note the
uncached column is itself content-sensitive (7.75 on noise vs 10.73 on
fruit.png), so the pipeline was never truly content-neutral either.
What it means for the next step. The keyed table is the design to carry — single-entry loses on most content and wins big only on near-solid fills. Content selection matters more than any tuning: photographs do not clear break-even, graphic and flat content clears it comfortably. And the heavier the pipeline, the better the cache looks, which makes CMYK destinations the recommendation and RGB→RGB the worst case.
For the kernel port, see "What this implies for a kernel port" above — the dispatch cost that dominates here vanishes in an unrolled loop, so the port is more attractive than this document first concluded, especially for the 4D/CMYK kernel. A POC on one kernel would settle it; a corpus of screenshots, halftones and print-origin scans would settle whether real workloads clear break-even.
src/cache.js, attached to Transform.prototype like stages.js and
interp.js. The 2026-08-17 default was 0. That changed — see
As built 2026-08-22.
Read counters with getPixelCacheStats() →
{enabled, slots, hits, misses, lookups, hitRate}; also
resetPixelCacheStats() and clearPixelCache().
Two stages are injected after init() (_applyPixelCache), so the
kernel can promote 'auto' to 1. The walk in transform() gained a
third arm (if pipelineDebug … else if cache … else …) that reads
stage.step.
Three things building it changed about the design notes:
-
The boundary can't be located from encodings or options. First attempt scanned encodings — but
stage_device_to_intlabels its outputdevice, and encoding3is bothLabD50(an object) andPCSXYZ(an array). Second attempt used options (convertInputOutput ? 1 : 0) — but a Lab input profile's stage 0 emits a{L,a,b}object, so position 1 holds nothing cacheable, and it scored zero hits everywhere. The position also depends on what the optimiser did, which no option records. Resolved by walking one probe colour from_buildValidationInput()at build time and taking the first position holding a numeric array of the right length — the same techniquevalidatePipeline()already uses. The store position uses a marker to the first output-conversion stage. LUT +stage_device_to_intis a legal fusion — if the marker is gone, inject stores at the end of the pipeline and the hit path copies. -
A fixed quantiser in the hash is a bug, not a detail. The first hash used
(value * 65536) | 0, assuming device floats in 0..1. The boundary often holds raw integers instead (0..255 / 0..65535), where that multiply pushes all the entropy out of range and collapses distinct colours into one slot — an interleaved 3-colour test scored 9 hits instead of 27. There are now two build-time variants:stage_pixelCache_keyedInt(oneimulon the value) andstage_pixelCache_keyed(hashes the double's raw bits, scale-free).dataFormatcannot choose between them —*sRGB- int8 leaves raw ints at the boundary while
*Lab+ int8 leaves floats — so the variant is detected from the probe colour. Misdetection is harmless: the hash only picks a bucket and the stored key is still compared exactly, so a wrong guess costs hit rate, never correctness.
- int8 leaves raw ints at the boundary while
-
transformArrayneeded a separate implementation. Its per-pixel walks increment blindly, so they would re-run the maths on a value a hit had already resolved. Acache active?test inside those loops measured ~2.5% on the uncached path, so the cached case now routes to a generic_transformArrayCached()and the unrolled loops are byte-identical to before. (That 2.5% later proved to be mostly measurement noise on a loaded box — a controlled A/B put it near 1% — but the restructure makes "cache-off costs nothing" structural rather than something to re-measure.)
Declines rather than misbehaves when it cannot guarantee
correctness: pipelineDebug on (a jump would fabricate a history that
never ran), custom stages present (a hit would skip their side
effects), or no numeric-array position before the store. A fused
output conversion no longer declines — store goes at the end.
Everything measured above is 3- and 4-channel, because those were the only profiles that existed. With the synthetic set (see SyntheticProfiles.md) the wide inputs can be measured too, and the answer is not a scaled-up version of the RGB one.
KernelND (7–15) declines the LUT — an A2B bake is grid^n cells — so
every pixel walks the pipeline at ~0.8 MPx/s rather than tens. A miss
therefore costs roughly fifty times what it costs in RGB, and the
whole economic argument moves with it. 5/6 now bake and run int8 WASM;
in-kernel inject (paired exports) is
__tests__/pixel_cache_wasm_5d_6d.tests.js. The accuracy-path table
still applies when buildLut is off or the kernel is KernelND.
Reproduce with:
node bench/pixel_cache/nchannel_bench.js4096 slots, int8, N → 3 channels, 60k px. off and on are MPx/s:
| in | content | distinct | off | on | gain | hit% |
|---|---|---|---|---|---|---|
| 4 | noise | 60000 | 6.38 | 4.86 | 0.76× | 0% |
| 4 | flat, 256 colours | 251 | 7.50 | 26.42 | 3.52× | 99% |
| 4 | flat, 16 colours | 16 | 7.74 | 28.83 | 3.72× | 100% |
| 5 | noise | 60000 | 3.06 | 2.66 | 0.87× | 0% |
| 5 | flat, 256 colours | 252 | 3.40 | 22.84 | 6.72× | 99% |
| 5 | flat, 16 colours | 16 | 3.45 | 26.07 | 7.55× | 100% |
| 8 | noise | 60000 | 0.75 | 0.70 | 0.94× | 0% |
| 8 | flat, 256 colours | 256 | 0.81 | 17.11 | 21.23× | 99% |
| 8 | flat, 16 colours | 16 | 0.83 | 21.50 | 25.79× | 100% |
| 12 | noise | 60000 | 0.68 | 0.63 | 0.93× | 0% |
| 12 | flat, 256 colours | 256 | 0.76 | 13.57 | 17.87× | 99% |
| 12 | flat, 16 colours | 16 | 0.88 | 18.37 | 20.82× | 100% |
(The 4-channel row runs with buildLut: false, because the cache lives on the
accuracy path and a CMYK transform would otherwise take the CLUT and never
reach it. It is the comparator, not the CMYK recommendation.)
The 25× is not the finding — 16 distinct colours is not a workload. The finding
is underneath it. Two ends of the same measurement give the whole model: the
noise rows are a pure miss (off without a cache, on with the lookup and
store the cache adds), and the flat-16 on row is a pure hit. Solve for the
rate at which they cancel:
h·tHit + (1−h)·tMiss = tOff
| input width | miss tax | hit vs the maths | break-even hit rate |
|---|---|---|---|
| 4 | 31% | 5× | ~29% |
| 5 | 15% | 9× | ~14% |
| 8 | 6% | 29× | ~6% |
| 12 | 8% | 27× | ~8% |
Two independent runs put 4 channels at 29–40%, 5 at 14–20%, and 8–12 at 6–12%, so read these as a shape rather than as constants. The shape is unambiguous: break-even falls by roughly a factor of four once the LUT is declined.
Both terms move in the same direction, which is why the effect is that large. A miss on the per-pixel path is so expensive that the lookup added to it barely registers — a 31% tax in CMYK becomes 6% at 8 channels — while a hit skips a correspondingly larger amount of work, 5× the maths in CMYK against 29× at 8 channels. Cheaper to be wrong, and far more valuable to be right.
The practical consequence is a threshold change, not a speed change.
docs/Transform.md says break-even is around 40% for RGB, "which flat graphic
content clears easily and photographs generally do not" — photographs measured
3–41%. Against a 6% bar, most of that range clears. Content this cache was
correctly judged not worth enabling for is worth enabling for above 4 channels,
and the crossover sits between 4 and 5, exactly where KernelND starts
declining the LUT. That is not a coincidence: it is the same cause.
What has not changed is that pure noise never pays, at any width. The cost shrinks with width — 0.76× at 4 channels, 0.94× at 8 — but it stays a cost. This cache rewards reuse; it cannot manufacture it.
All three produced convincing numbers before being caught, and the first version of this section was published with the third one in it.
pixelCache: true is ONE slot. It resolves to 1, not "on with a sensible
table" — and 'auto' (when a kernel promotes it) is also 1. That is
deliberate for unknown content; see Why auto is 1.
Pass a count when you know the work is a palette: pixelCache: 4096.
Reading a 13% hit rate on content with eight distinct colours is the
symptom, because a single entry hits 1-in-8.
An LCG's low byte has a short period. Generating "noise" as s & 0xff
produces content that repeats far more than random, which flatters the cache
into reporting near-perfect hits on supposedly-unique pixels. Take the high
bits — (seed >>> 16) & 0xff, which is what cache_bench.js already did — and
print the distinct-colour count to check what was actually generated.
Timing N passes over one buffer warms the cache between them. This is the
one that produced a wrong published result, and it is the least visible: build
one Transform, run the same image through it three times, take the best. Pass
2 finds pass 1's entries resident, so unique content arrives with a 4096-entry
head start and "best of three" selects the most warmed pass. It read as 17%
reuse on data that has none, and turned a 0.94× cost at 8 channels into a
reported 1.20× gain — which then supported a conclusion about photographic
content that the data did not actually contain. The harness now builds a fresh
Transform per timed pass and warms the JIT on a throwaway one.
The correct conclusion survived; the number under it did not. The break-even framing above is the version that holds, and it is stronger — it says why the width matters instead of asserting that it does.
__tests__/pixelcache.tests.js covers 35 input×output width combinations from
1 to 15 channels, both depths, asserting cached output is byte-identical to
uncached — plus that the cache genuinely engages at those widths rather than
quietly declining, and that where it does decline (an identity pair) the
conversion is still correct. Also: 'auto' — 4/5/6 promote to 1 and inject;
3CLR leaves it; array() still names the kernel.
Transform does not look at inputChannels. The kernel that won init()
does.
| hint | what Transform does | what the kernel may do |
|---|---|---|
'auto' (default) |
ignore | 4/5/6 change it to 1; 3D / 1D / 2D / identity / ND / matrix-shaper leave it |
0 / false |
off | leave it |
1 / true |
single-entry | leave it, except a 3-D matrix pair forces 0 and still yields the shaper |
16, 32, 256… |
that many slots (power of two) | leave it |
After init(), _applyPixelCache() injects only if the value is then a
number > 0. pixelCacheUsed is what landed (0 if inject declined).
kernelInfo().cache is 'not-supported' | 'off' | 1 | N.
opts.pixelCache and opts.pixelCacheActive (forced number, not auto)
are named facts on _kernelOpts(). A kernel that throws from init()
is ignored; auto stays auto.
The accuracy-path cache is two pipeline stages. transform() walks
them. array() on a bound LUT kernel (enableForArrays) does not
— it calls kernel.array(), which is the WASM / JS int path. Auto on
4/5/6 therefore memos single colours and does not steal the image
path. That is why mpx_summary LUT cells stay on kernel3D /
kernel4D.
If there is no kernel batch path (no LUT, not claimed, objects),
array() uses _transformArrayCached when the stages are in.
Two products got measured. They disagree about tables.
In-kernel WASM (paired exports, run_paired_*.js) — single-entry
vs a 2 / 8 / 16 / 32 / 256-slot table, same bits out:
| path | photo+5% noise tax, single | photo+5% noise tax, table | solid, single |
|---|---|---|---|
| 3D SIMD int8 | 2% | 7–9% | 3.1× |
| 4D SIMD int8 | 3% | 10% | 5.2× |
| 4D scalar int8 | 1% | 7% | 6.6× |
| 5D scalar | 16% | worse | 13× |
| 6D scalar | 1% | ~1% | 39× |
int16 is the same policy (slightly smaller solid win, same photo
band). A table is slower than single on flats and a worse miss
tax on photo/noise. It would only win on a small palette of
non-adjacent repeats — you cannot see that from inputChannels.
So the in-kernel product is still 1. Table snippets stay in the
POC builder; they are not in the shipped *.wasm.js.
TODO — two-register double, not a 2-slot array. We never injected “last two keys, last two outputs” as two locals. That is still just (1) with a second compare. An array starts at the first hashed slot in memory; whether 2 or 4096, it is a different tax. Build the double only if dither/ABAB is a real workload; do not infer it from the table arms.
Accuracy path — table size is almost free (4096 no slower than
32) and a keyed table does catch interleaved palettes that a
single entry misses (A,B,C,A,B,C… is 0 hits on one slot, 27 on 32).
That is why pixelCache: 256 is still the right explicit choice
when you know the work is a poster.
Auto does not know the work. It picks last-pixel, which is what
unknown content actually is: solids, logos, and runs. A 2+ table
needs a size, and the size that is free to have is not free to
guess — true resolving to 1 already surprised people who
thought it meant "on with a sensible table". Auto meaning 256 would
quietly put a 256-slot walk on every 4CLR transform() and still
lose on a solid to the one-slot form.
So: auto → 1. Want 2+? Pass the count.
The one number that is not "1–3% for a 4–6× solid win" is 5D
photo with 5 % noise added at 0.84× on the in-kernel
single-entry export. Auto
still turns 5D on: 5CLR input is almost never a grainy photograph.
Leave pixelCache: 0 if that array() cell matters.
One hint, two implementations:
| path | who | 'auto' |
0 |
report |
|---|---|---|---|---|
transform() |
src/cache.js stages |
4/5/6 inject 1; 3D leaves | off | pixelCacheUsed |
array() WASM 3–6 |
interp_*_cached |
bind cached export | verbatim export | kernelInfo().cache === 1 |
Matrix-shaper / identity / 1D / 2D / ND / JS fallback: 'not-supported'
or 'off'. No second option name — the kernel that can bind does;
the one that cannot declines. Hash tables are not shipped.
Does the accuracy-path hypothesis hold?Yes above 4 input channels, and on flats at 3/4. Photographs on 3D still do not.- Does the run-scan restructure actually beat the memo, or does V8 handle the spill better than expected?
- Is dithered continuous-tone input a real workload? Normally you transform before screening, and post-screen 1-bit data never reaches a colour transform — if so, option (2) and (3) both lose their justification and (1) is the whole story.
- Interaction with
preserveAlphaand the identity_kernelCopypath — both already skip work; confirm no double-counting. - Two-register double — last two
u32keys, two compares, no array. Never built. See option 2.