Skip to content

Use the vector index for related articles - #238

Merged
ssavutu merged 1 commit into
mainfrom
perf/related-articles-vector-index
Sep 27, 2026
Merged

ssavutu merged 1 commit into
mainfrom
perf/related-articles-vector-index

Conversation

@ssavutu

@ssavutu ssavutu commented Sep 27, 2026

Copy link
Copy Markdown
Member

Why

GetRelatedArticlesBySlug runs on every article view. It joined articles and both sides of article_embeddings before ORDER BY VEC_DISTANCE_EUCLIDEAN(...), which disqualifies the HNSW index (the same trap buildVectorNeighbourQuery documents for search). Each view scanned all ~10k vectors and filesorted them. That's ~20ms warm and >1s under load, and it accounts for 358 of 370 entries in the replica's slow log.

Change

  • Look up the source article's id (no stored vector → empty list, as before).
  • Rank with the same derived-table neighbour query as search, via a new buildRelatedNeighbourQuery. It reads the source vector through an uncorrelated subquery. Passing it back as text isn't an option because VEC_ToText is lossy: 0 of 10,097 production vectors round-trip byte for byte.
  • Exclude the source article, with one extra row of over-fetch for it.
  • Load the rows with LoadArticlesByIDsInOrder, then authors as before.

Verified

  • On production data (replica, read-only): EXPLAIN uses the embedding vector key, and the top 3 are identical to the old query for sampled slugs. Under 10ms.
  • New shape test, plus integration tests for ordering, self-exclusion, non-live filtering, and the no-vector case.
  • go test -p 1 ./internal/... passes against MariaDB 11.8.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Q39NMqRSaA8bY8FTaB6Uk5

GetRelatedArticlesBySlug ranked candidates in one statement that joined
articles and both sides of article_embeddings before ORDER BY
VEC_DISTANCE, which disqualifies the HNSW index. Every article view
scanned all ~10k stored vectors plus a filesort: 20ms warm, over 1s
under load, and 358 of the 370 entries in the replica's slow log.

It now looks up the source article's id, then goes through the same
derived-table neighbour query as vector search, reading the source
vector by subquery (VEC_ToText is lossy, so passing it back as text
would shift distances) and excluding the source from its own list.
Checked on production data: the plan uses the vector key, the top 3
are identical, and the query takes under 10ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q39NMqRSaA8bY8FTaB6Uk5
@ssavutu
ssavutu merged commit 897b010 into main Sep 27, 2026
7 checks passed
@ssavutu
ssavutu deleted the perf/related-articles-vector-index branch September 27, 2026 08:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant