Skip to content

fix(fts): snapshot FTS aux info in bind data (Fixes ladybugdb/ladybug#1105) - #94

Merged
adsharma merged 2 commits into
mainfrom
fix-1105-fts-bind-dangling-entry
Oct 7, 2026
Merged

adsharma merged 2 commits into
mainfrom
fix-1105-fts-bind-dangling-entry

Conversation

@adsharma

@adsharma adsharma commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Second parameterized CALL QUERY_FTS_INDEX on the same connection segfaults (LadybugDB/ladybug#1105, follow-up to LadybugDB/ladybug#1082 / LadybugDB/ladybug#1088).

Root cause: QueryFTSBindData held a const IndexCatalogEntry reference into the versioned catalog. Bind data outlives the bind transaction via the prepared-plan cache, so the 2nd execution dereferences a dangling entry.

Fix: store an owned FTSIndexAuxInfo value copy; read config from auxInfo in getQueryTerms/bindFunc/writer.

Minimal local testing per policy - CI owns build + e2e (extension rebuild + 4-mode repro from the issue).

…#1105)

QueryFTSBindData held a reference into the versioned catalog,
which dangles when the prepared plan is reused on a second
parameterized execution on the same connection, segfaulting
(Refs LadybugDB/ladybug#1082, Refs LadybugDB/ladybug#1088).
Store an owned copy of FTSIndexAuxInfo instead.
@idanasis

idanasis commented Oct 6, 2026

Copy link
Copy Markdown

Tested this PR's CI build (lbug-extensions-linux-x86_64 from run 37218957218, fts/libfts.lbug_extension, loaded with LOAD EXTENSION '<path>') against ladybug 0.21.2 on Linux x86_64 (python:3.11-slim, no network, so the official extension cannot be fetched). The control is the official INSTALL fts extension in the same image.

The #1105 crash is fixed:

same connection, $-parameterized QUERY_FTS_INDEX official fts (0.21.2) this PR
sync, 3 calls SIGSEGV on 2nd call (exit 139) OK, correct results
AsyncConnection, 3 calls SIGSEGV on 2nd call OK
50 calls cycling 5 terms n/a 0 mismatches
insert a doc between calls, then query its term n/a finds it
40 concurrent async calls (max_concurrent_queries=4) n/a 0 mismatches

Remaining gap: the cached parameterized plan is not invalidated when the FTS index is dropped or recreated. Same connection:

Q = "CALL QUERY_FTS_INDEX('Doc', 'doc_fts', $q) RETURN node.id AS id ORDER BY id"
c.execute(Q, {"q": "alpha"})                      # ['0']
c.execute("CALL DROP_FTS_INDEX('Doc', 'doc_fts')")
c.execute(Q, {"q": "alpha"})                      # this PR: still returns ['0'] (index is gone)
c.execute("CALL CREATE_FTS_INDEX('Doc', 'doc_fts', ['text'])")
c.execute(Q, {"q": "alpha"})                      # this PR: RuntimeError: unordered_map::at

With the literal form of the same query (QUERY_FTS_INDEX('Doc', 'doc_fts', 'alpha')), both the official extension and this PR behave correctly:

  • after the drop: Binder exception: Table Doc doesn't have an index with name doc_fts;
  • after the recreate: ['0'].

A new Connection is also correct. So it looks specific to the reused prepared plan now holding its own FTSIndexAuxInfo copy across catalog changes. Before this PR that path segfaulted on the second call, so it was never reachable.

The prepared-plan cache reuses QueryFTSBindData across executions on
the same connection, so the snapshotted aux info / backing-table graph
entry / index stats went stale after DROP_FTS_INDEX (stale results) and
broke after CREATE_FTS_INDEX (dropped table IDs).

Store the queried table/index names and re-resolve them against the
catalog on every execution (refreshFromCatalog). A missing index throws
the same BinderException as the literal path, and a recreated index
refreshes the snapshot. The cached physical plan also reuses the shared
graph, so rebuild it when the backing table IDs changed.
@adsharma
adsharma merged commit 2b9d5f4 into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants