Repository navigation
fix(storage): harden checkpoint crash recovery - #1099
Conversation
|
Why Both fd199c1 switches the assertion to One pre-existing limitation I ran into while validating this
|
|
This looks great! Ready to merge. A few thoughts I'm jotting down for future reference:
|
e8253de to
7111ecd
Compare
|
On the two notes: Parent database ID: nothing in the recovery path keys a child graph to its parent's database ID. Graph WAL headers verify against the graph's own persisted ID ( Downgrade stance: that matches the intent. The documented gate is a successful |
Description
A crash inside a checkpoint could leave on-disk state that recovery
misinterpreted, and several failure windows could silently lose or corrupt
data. This change makes the WAL checkpoint record the single commit point of
the protocol: everything recovery needs (main, graph, and partition-child
shadow files, plus each child's database header and page manager) is flushed
and fsynced before the commit record is written, and once that record is
durable, a failure anywhere makes the database refuse further writes until
restart instead of rolling back over a committed checkpoint. Shadow replay is
bounded to the page extent committed in the checkpoint header, locates the
last database-header record and validates its database ID before applying
anything, and rejects any record that targets a page outside that extent.
Shadow files are stamped with the ID of the database running the checkpoint,
using a bundle sentinel for the new format, so a graph's or partition child's
shadow names the parent that owns its pending bundle. A standalone open of a
file whose shadow belongs to another database's bundle refuses to decide that
bundle's fate instead of discarding the shadow. A graph whose file is missing
while its shadow is present under a committed checkpoint fails loudly instead
of being skipped, and a graph whose active and frozen WALs both carry
checkpoint markers fails instead of half-recovering. The WAL CHECKPOINT
record is versioned: this build writes a bundle-format marker, parses records
without it as legacy v0, and rejects unknown future versions. WAL rotation
failures poison the database instead of continuing on a half-rotated WAL,
retired WALs are removed only after the checkpoint's shadow state is fully
applied, and directory renames and removals are fsynced.
Two windows that previously corrupted data on retry are fixed. A checkpoint
retried after a failed attempt used to skip hash-index header pages that the
failed attempt had already advanced past, which corrupted the header chain
and made recovery fail with "disk array header page 0"; rollback now restores
the staged header-page count so the retry rewrites the chain. And databases
written by the legacy checkpoint protocol, which applied and removed each
partition child's shadow before retiring the main WAL, used to fail recovery
when a child shadow was legitimately absent; legacy markers now treat a
missing child shadow as already applied, while bundle-format checkpoints
still require every shadow. Each window has a regression test that fails on
the old behavior.
docs/checkpoint_recovery.md documents the resulting protocol: the commit
order, what each marker means, and what recovery does for each crash point.
Types of changes
Checklist
Not needed: the database file layout is unchanged. The additions are
confined to transient recovery artifacts and are self-describing: a
trailing
ownerDatabaseIDin the shadow-file header (zero reads as thefile's own bundle) and a trailing versioned field in the WAL CHECKPOINT
record (absent reads as legacy v0), so existing databases and older shadow
files open unchanged.