fix(storage): don't zero CSR header and length-chunk buffers - #1114
Merged
adsharma merged 1 commit intoOct 5, 2026
Merged
Conversation
Three createColumnChunkData calls passed a trailing false meant as initializeToZero, but the factory's sixth parameter is hasNullData, so the UINT64 buffers were calloc'ed: the CSR header of every rel scan state (per worker thread per SCAN_REL_TABLE), and the length chunks in CountRelTable and RelTable::getDegreeEntries. Pass both flags explicitly. Nothing reads these buffers before writing them. Fixes LadybugDB#1110. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
|
Thank you! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This work was produced with the help of language models.
Fixes #1110.
Every
ScanRelTableworker zeroed a 2 MiB CSR header before it looked for a morsel. ThreeColumnChunkFactory::createColumnChunkDatacalls passed a trailingfalsemeant asinitializeToZero, but the signature ends inbool hasNullData = true, bool initializeToZero = true, so it bound tohasNullDataandmallocBuffercalloced theUINT64buffers. The sites: theInMemChunkedCSRHeaderconstructor (every rel scan state owns one atNODE_GROUP_SIZEcapacity, one per worker thread), and the length chunks inCountRelTable(its comment already said/*initializeToZero*/) andRelTable::getDegreeEntries.Fix. Pass both flags explicitly:
false /*hasNullData*/, false /*initializeToZero*/.Why it is safe. Nothing reads these buffers before writing them, every header getter clamps to
numValues, and flushing coversnumValuesentries, so an unwritten tail reaches neither stats nor disk. A build that fills all three buffers with0xABright after allocation passes the fullctestsuite.Measurements. The #1110 reproducer (100
Tagnodes, onehasTypeeach), median of three runs of 200 timed executions,mainat 305d80e before → after:MATCH (t:Tag)-[:hasType]->(k:Class) WHERE t.name = 'tag7' RETURN count(*)The repo's LDBC SNB SF1 suite (54 queries): geomean 0.92x at 16 threads and 0.93x at 1 thread; the 23 queries with a relationship pattern are at 0.84x, the 31 node-only queries at 0.98x and 1.00x. Every query returned identical output on both builds.
Tests.
EmptyBufferManagerTest.CSRScanStateHeaderIsTwoPlainUInt64Buffers(test/storage/buffer_manager_test.cpp) pins the header a rel scan state carries: twoUINT64buffers ofNODE_GROUP_SIZEcapacity, no null chunk, the buffer pool charged exactly their size. It also passes on the unpatched tree:mallocBufferaccounts the bytes before it choosescallocormalloc, so zeroing is invisible to every counter the engine exposes; the measurements above are the evidence it is gone.ctest(make test-build-release): 0 of 2861 failed (23 skipped), 258 disabled by the suite. clang-format 18.1.8 reports no changes.Kuzu history, write-before-read sites, full table and methodology
kuzu 0.11.3 built the header through
ColumnChunk(…, residencyState, false), whose last parameter isinitializeToZero. The Segmentation refactor (kuzudb/kuzu#5950) moved the call to the factory, where the samefalselands onhasNullData; kuzumasterstill has that call.Where each buffer is written before it is read: the scan path resets
numValuesto 0 and appends the on-disk header throughscanCommitted; checkpoint writes withcopyFromorpopulateCSRLengthInMemOnly+populateStartCSROffsetsFromLength;COPYstarts with thestd::fillinpopulateCSRLengthsInternal; the two length chunks are filled bycsrLengthColumn->scanbefore they are summed.getMetadataToFlushandflushBuffercovernumValuesentries.Reproducer setup: against the C API, both builds
make benchmark(Release, gcc 13.3), idle 16-core Ryzen 9 6900HS,CHECKPOINTbefore timing,max_num_threads = 16,lbug_connection_set_max_num_thread_for_exec(N). Rows not shown above:MATCH (t:Tag)-[:hasType]->(k:Class) WHERE t.name = 'tag7' RETURN count(*)MATCH (t:Tag)-[:hasType]->(k:Class) RETURN count(*)(plans asCOUNT_REL_TABLE)In a long-running process the freed buffers are reused, so the cost shows up as CPU spent zeroing rather than as page faults; the Python script in #1110 starts a fresh process and goes from 5492 to 19 minor faults per execution at 16 threads (0.21.1, header fix applied).
Benchmark suite:
benchmark/queries/ldbc-sf100on LDBC SNB SF1, 1 warm-up and 5 runs per query, six rounds alternating which build runs first, median of the 30 runs. Largest gains at 16 threads:join/SelectiveTwoHopJoin3.58 → 1.37 ms,join/q311.71 → 0.72 ms,ldbc_snb_ic/q366.15 → 3.99 ms (join/q311.17 → 0.28 ms at 1 thread). Individual node-only ratios span 0.82x–1.08x at 16 threads, the noise band for single-query ratios here.