Skip to content

fix(wal): tag WAL records with their owning catalog and route replay per owner - #1113

Open
aikins01 wants to merge 15 commits into
LadybugDB:mainfrom
aikins01:fix/wal-owner-tagging
Open

aikins01 wants to merge 15 commits into
LadybugDB:mainfrom
aikins01:fix/wal-owner-tagging

Conversation

@aikins01

@aikins01 aikins01 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Description

Every WAL record is tagged with the name of the catalog that owns the object it
describes, and recovery resolves each record against that owner instead of the
ambient main catalog. Object IDs are per-catalog, so before this change a
replayed graph-owned record (a sequence advance, a node insert) resolved its
OIDs through main's catalog and either asserted or landed on the wrong object.
Ownership is captured at record-creation time from the affected table or entry,
so transactions that cross graphs via USE GRAPH are handled per record, and
BEGIN/COMMIT records stay scope-neutral.

The wire format appends a length-framed ownerCatalogName field to every
record, which an older reader skips harmlessly. The empty string resolves
through the ambient replay scope (main for main-WAL records, the owning
graph's catalog while a graph WAL replays), and the field is absent in legacy
WALs, which resolve to main exactly as before. WAL::CHECKPOINT_BUNDLE_FORMAT_VERSION
is bumped so older builds refuse a new checkpoint bundle instead of silently
skipping the tags and misrouting its records. Replay wraps each record in a
recovery-only owner scope that the ambient Catalog::Get / StorageManager::Get
chain consumes ahead of the session's default graph; a graph whose create-graph
record is earlier in the same WAL is materialized on demand before its data
records replay. Local storage is keyed by (owner, tableID) so one replayed
transaction that touched two graphs cannot conflate their local state.

Commit tracks the set of changed catalogs and bumps each one's version through
the database manager's registry lock, so a checkpoint never skips serializing a
graph catalog changed by replayed DDL, and a catalog destroyed by a DROP GRAPH (in the same transaction or on a concurrent connection) is never
touched after destruction.

Records whose owning graph was dropped after the record committed are skipped
during replay: DROP GRAPH reclaims the graph's files immediately, while the
committed records stay in the main WAL until the next checkpoint, and the
graph can no longer be materialized. Throwing there would wedge recovery
permanently, since the committed prefix can never replay past the record and
the WAL can never be retired.

Replayed DDL recreates entries through CatalogSet::createEntry, which
regenerates per-catalog OIDs, so a recorded ID can name a different object at
replay time, most commonly because ANY-graph lazy materialization recreates
the _nodes/_edges infrastructure entries before retained records replay,
shifting every later ID in the recovered catalog. Replay therefore records a
recorded-to-replayed entry mapping for created tables, sequences, and indexes,
and drop/alter records resolve their entry IDs through those mappings before
falling back to the raw ID. Without it, a standalone session's recorded
DROP INDEX resolved its raw recorded OID against the recovered catalog and
could drop a parent-created index that happened to occupy that OID after
replay. The mappings are scoped to the replayer instance and never shared
across recording sessions, and graph-WAL replay uses its own replayer
(replayPendingGraphWALs), so main-WAL recorded IDs can never be translated
through graph-WAL mappings.

Depends on #1099 (its versioned checkpoint record and bundle-format gate) and
should merge after it.

Known limitations

  • UPDATE_SEQUENCE replays the recorded sequence ID. Sequence records appear
    in commit order, so the recorded ID keeps naming the same sequence except
    across a drop-and-recreate of sequences captured in one retained WAL.
  • Rel-group and partition-parent ALTER records still carry raw physical table
    IDs, and a raw ID can land on a node table at replay and wedge recovery in a
    cast failure. This predates the change: owner tagging decides which catalog
    a record replays against, not how ALTER resolves physical IDs.
  • Reopening a database a second time while a graph's WAL holds records from
    two ID namespaces (one recorded before an earlier recovery replayed the
    graph's WAL, one after) can silently reinterpret object IDs across tables.
  • Owner resolution in WALReplayer::replayWALRecord walks the registry linearly
    with case normalization (hasGraph twice, getGraphCatalog once), so
    replaying R relationship inserts across G registered graphs costs Θ(R×G)
    name comparisons. Recovery is offline, and modest graph counts keep this
    cheap.

Types of changes

  • Bug fix
  • New feature
  • Breaking change
  • Documentation Update

…owner

Graph table commits log to the single main WAL, but replay resolved
table OIDs through the ambient catalog, so committed graph data resolved
against the wrong catalog's OID space at recovery. Records now carry
their owning catalog's name on the wire (checkpoint bundle format
version 2), catalog entries back-point to their owning catalog, replay
resolves each record through its owner, and local storage keys staged
tables by owner so cross-graph transactions commit each table to its
own graph.

Main's catalog carries no storage manager, so owner-aware storage
resolution falls back to the database storage manager for main. A
tagged record whose graph was created in the same WAL (and is therefore
not loaded yet during main-WAL replay) materializes it on demand, and
materialized ANY graphs whose catalogs were never checkpointed get
their infra tables recreated so OID spaces line up.

Checkpoint recovery recognizes every persisted bundle-format version
1..CHECKPOINT_BUNDLE_FORMAT_VERSION as a bundle: a version-1 marker
written by an earlier build must still apply graph shadows instead of
taking the pre-bundle legacy path that discards them.
…atalogs

Replay graph-owned uncheckpointed records into the owner's catalog,
cross-graph manual transactions into each owner, and a graph's DDL plus
first row committed to a single WAL. Assert row content per graph so
cross-graph swaps are caught, including rel inserts staged for two
graphs in one manual transaction. A checkpoint bundle persisted by an
earlier bundle-format version must recover through the bundle path and
keep applying graph shadows.
Keep the loadGraph helper in loadGraphsFromCatalog and port main's
graph-catalog function fallback into it; keep both sides' includes in
checkpoint_test.cpp.
@aikins01

aikins01 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

The pinned vector extension (extension/vector/src/index/hnsw_graph.cpp:713-716) calls the table-ID-keyed Transaction::isUnCommitted/getLocalRowIdx overloads. Since this PR makes table IDs per-catalog, a table-ID-only lookup can't identify the owning catalog in general, so the core API became table-scoped and the pinned extension no longer compiled.

11515d628 restores the two overloads as a compatibility shim: they consult the transaction's local storage and match a table ID across owning catalogs, answering only when exactly one owner holds it (and "no local table" when none or several do). This keeps the pinned submodule green without changing its behavior.

The proper extension-side fix is up as LadybugDB/extensions#95 (pass the node table directly); once it lands and the submodule pointer is bumped, the shims can be removed.

@aikins01

aikins01 commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Review follow-ups on the last three commits:

  • 483500847: named UPDATE_SEQUENCE records travel under their own record type, so the sequence name can never be misread as the legacy [sequenceID][kCount] layout. The legacy type still decodes from binaries written before this change.
  • a06a44dba: the graph registry lookups now key their recursion by manager and re-check liveness per entry.
  • 8fcc67cdd: the legacy-format regressions no longer build payloads through the current writer. The deserialization test hand-writes the pre-named-format bytes, so a matching incompatible change to both writer and reader can't still pass as legacy decoding. The new replay test appends a crafted legacy transaction whose recorded entry ID shifts across replay, pinning the empty-name fallback to the create-replay ID map: replay must advance the crafted sequence and leave the table's serial untouched.

Full transaction suite passes on the branch tip (172 tests); the two sequence tests also pass with checksums disabled.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant