Conversation
…owner Graph table commits log to the single main WAL, but replay resolved table OIDs through the ambient catalog, so committed graph data resolved against the wrong catalog's OID space at recovery. Records now carry their owning catalog's name on the wire (checkpoint bundle format version 2), catalog entries back-point to their owning catalog, replay resolves each record through its owner, and local storage keys staged tables by owner so cross-graph transactions commit each table to its own graph. Main's catalog carries no storage manager, so owner-aware storage resolution falls back to the database storage manager for main. A tagged record whose graph was created in the same WAL (and is therefore not loaded yet during main-WAL replay) materializes it on demand, and materialized ANY graphs whose catalogs were never checkpointed get their infra tables recreated so OID spaces line up. Checkpoint recovery recognizes every persisted bundle-format version 1..CHECKPOINT_BUNDLE_FORMAT_VERSION as a bundle: a version-1 marker written by an earlier build must still apply graph shadows instead of taking the pre-bundle legacy path that discards them.
…atalogs Replay graph-owned uncheckpointed records into the owner's catalog, cross-graph manual transactions into each owner, and a graph's DDL plus first row committed to a single WAL. Assert row content per graph so cross-graph swaps are caught, including rel inserts staged for two graphs in one manual transaction. A checkpoint bundle persisted by an earlier bundle-format version must recover through the bundle path and keep applying graph shadows.
Keep the loadGraph helper in loadGraphsFromCatalog and port main's graph-catalog function fallback into it; keep both sides' includes in checkpoint_test.cpp.
|
The pinned vector extension (
The proper extension-side fix is up as LadybugDB/extensions#95 (pass the node table directly); once it lands and the submodule pointer is bumped, the shims can be removed. |
|
Review follow-ups on the last three commits:
Full transaction suite passes on the branch tip (172 tests); the two sequence tests also pass with checksums disabled. |
Description
Every WAL record is tagged with the name of the catalog that owns the object it
describes, and recovery resolves each record against that owner instead of the
ambient main catalog. Object IDs are per-catalog, so before this change a
replayed graph-owned record (a sequence advance, a node insert) resolved its
OIDs through main's catalog and either asserted or landed on the wrong object.
Ownership is captured at record-creation time from the affected table or entry,
so transactions that cross graphs via
USE GRAPHare handled per record, andBEGIN/COMMITrecords stay scope-neutral.The wire format appends a length-framed
ownerCatalogNamefield to everyrecord, which an older reader skips harmlessly. The empty string resolves
through the ambient replay scope (main for main-WAL records, the owning
graph's catalog while a graph WAL replays), and the field is absent in legacy
WALs, which resolve to main exactly as before.
WAL::CHECKPOINT_BUNDLE_FORMAT_VERSIONis bumped so older builds refuse a new checkpoint bundle instead of silently
skipping the tags and misrouting its records. Replay wraps each record in a
recovery-only owner scope that the ambient
Catalog::Get/StorageManager::Getchain consumes ahead of the session's default graph; a graph whose create-graph
record is earlier in the same WAL is materialized on demand before its data
records replay. Local storage is keyed by (owner, tableID) so one replayed
transaction that touched two graphs cannot conflate their local state.
Commit tracks the set of changed catalogs and bumps each one's version through
the database manager's registry lock, so a checkpoint never skips serializing a
graph catalog changed by replayed DDL, and a catalog destroyed by a
DROP GRAPH(in the same transaction or on a concurrent connection) is nevertouched after destruction.
Records whose owning graph was dropped after the record committed are skipped
during replay:
DROP GRAPHreclaims the graph's files immediately, while thecommitted records stay in the main WAL until the next checkpoint, and the
graph can no longer be materialized. Throwing there would wedge recovery
permanently, since the committed prefix can never replay past the record and
the WAL can never be retired.
Replayed DDL recreates entries through
CatalogSet::createEntry, whichregenerates per-catalog OIDs, so a recorded ID can name a different object at
replay time, most commonly because ANY-graph lazy materialization recreates
the
_nodes/_edgesinfrastructure entries before retained records replay,shifting every later ID in the recovered catalog. Replay therefore records a
recorded-to-replayed entry mapping for created tables, sequences, and indexes,
and drop/alter records resolve their entry IDs through those mappings before
falling back to the raw ID. Without it, a standalone session's recorded
DROP INDEXresolved its raw recorded OID against the recovered catalog andcould drop a parent-created index that happened to occupy that OID after
replay. The mappings are scoped to the replayer instance and never shared
across recording sessions, and graph-WAL replay uses its own replayer
(
replayPendingGraphWALs), so main-WAL recorded IDs can never be translatedthrough graph-WAL mappings.
Depends on #1099 (its versioned checkpoint record and bundle-format gate) and
should merge after it.
Known limitations
UPDATE_SEQUENCEreplays the recorded sequence ID. Sequence records appearin commit order, so the recorded ID keeps naming the same sequence except
across a drop-and-recreate of sequences captured in one retained WAL.
IDs, and a raw ID can land on a node table at replay and wedge recovery in a
cast failure. This predates the change: owner tagging decides which catalog
a record replays against, not how ALTER resolves physical IDs.
two ID namespaces (one recorded before an earlier recovery replayed the
graph's WAL, one after) can silently reinterpret object IDs across tables.
WALReplayer::replayWALRecordwalks the registry linearlywith case normalization (
hasGraphtwice,getGraphCatalogonce), soreplaying R relationship inserts across G registered graphs costs Θ(R×G)
name comparisons. Recovery is offline, and modest graph counts keep this
cheap.
Types of changes