Repository navigation
fix(storage): refresh node group scan state after a checkpoint - #1094
Merged
adsharma merged 2 commits intoOct 4, 2026
Merged
Conversation
Read-only transactions are not blocked by a checkpoint, so a checkpoint can run between two vectors of the same node table scan (or between two lookups that share a scan state). NodeGroup::checkpoint merges all chunked groups into a single persistent one and rewrites its column chunk metadata, but the NodeGroupScanState of the running scan still holds the chunked group index and the per-segment metadata (page ranges, compression metadata, dictionary sizes) captured before the checkpoint. The scan then reads the checkpointed pages with stale metadata, or indexes past the end of the chunked group list. This returned corrupted values (strings assembled from the wrong dictionary offsets, wrong integers) without any error, and could crash when the scan was inside a committed in-memory chunked group that the checkpoint merged away. Count the checkpoints of a node group and record the count in the scan state when it is initialized. Scans and lookups, which already hold the chunked-groups lock, re-initialize the state against the checkpointed group when the counts differ. Row indices within a node group are stable across a checkpoint, so the scan position is kept.
Contributor
|
Does |
Contributor
Author
|
Yes. Rel scans snapshot their chunk state and CSR header once, and |
1 of 7 tasks
Contributor
Author
|
The CSR follow-up is #1116. A scan that is already running keeps the CSR state it started with. Checkpoint does not wait for it, and the node-table cursor refresh is not reused. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Bug. Read-only transactions are not blocked by a checkpoint, so a checkpoint can run between two vectors of the same node table scan (or between two lookups sharing a scan state).
NodeGroup::checkpointmerges all chunked groups into a single persistent one and rewrites its column chunk metadata, but the running scan'sNodeGroupScanStatestill holds the chunked group index and the per-segment metadata (page ranges, compression metadata, dictionary sizes) captured byinitializeScanStateForChunkedGroupbefore the checkpoint.NodeGroup::scanthen reads the checkpointed pages with stale metadata, or indexes past the end of the chunked group list.Symptoms: scans return corrupted values with no error (strings assembled from the wrong dictionary offsets, wrong integers), and a scan that was inside a committed in-memory chunked group crashes. Reproduces on main (v0.21.2) with 3 scanning connections, a delete/re-insert writer and a
CHECKPOINTloop at 8 threads: garbled strings in 3 of 8 two-minute runs.Fix. Count the checkpoints of a node group (
NodeGroup::numCheckpoints, protected by the chunked-groups lock) and record the count in the scan state when it is initialized.scan,scanInternal,lookupandlookupMultiple, which already hold that lock, re-initialize the state against the checkpointed group when the counts differ. Row indices within a node group are stable across a checkpoint, sonextRowToScanis kept. Cost: one integer comparison per scanned vector / lookup.Tests.
ScanAcrossCheckpointTest.CheckpointWhileScanningPersistentGroupand.CheckpointWhileScanningInMemoryGroupdrive aNodeTableScanStatein a read-only transaction and checkpoint between two vectors. Without the fix the first returns wrong values and the second segfaults; with it both pass. The stress loop above: 0 of 13 runs garbled with the fix.Verified locally: release
ctest2854/2855 (the one failure,dictionary_bug~orb383_relationship_projection_obfuscated.AnonymousParquetDeleteReload, also fails on main); ASan + runtime checks:transaction_test87/87, storage unit tests, e2etransaction~*:dml_node~*:dml_rel~*:storage*725/725.Not addressed here.
CSRNodeGroup::checkpointand may have the same exposure.Types of changes