Repository navigation
Use the vector index for related articles - #238
Merged
Merged
Conversation
GetRelatedArticlesBySlug ranked candidates in one statement that joined articles and both sides of article_embeddings before ORDER BY VEC_DISTANCE, which disqualifies the HNSW index. Every article view scanned all ~10k stored vectors plus a filesort: 20ms warm, over 1s under load, and 358 of the 370 entries in the replica's slow log. It now looks up the source article's id, then goes through the same derived-table neighbour query as vector search, reading the source vector by subquery (VEC_ToText is lossy, so passing it back as text would shift distances) and excluding the source from its own list. Checked on production data: the plan uses the vector key, the top 3 are identical, and the query takes under 10ms. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q39NMqRSaA8bY8FTaB6Uk5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
GetRelatedArticlesBySlugruns on every article view. It joinedarticlesand both sides ofarticle_embeddingsbeforeORDER BY VEC_DISTANCE_EUCLIDEAN(...), which disqualifies the HNSW index (the same trapbuildVectorNeighbourQuerydocuments for search). Each view scanned all ~10k vectors and filesorted them. That's ~20ms warm and >1s under load, and it accounts for 358 of 370 entries in the replica's slow log.Change
buildRelatedNeighbourQuery. It reads the source vector through an uncorrelated subquery. Passing it back as text isn't an option becauseVEC_ToTextis lossy: 0 of 10,097 production vectors round-trip byte for byte.LoadArticlesByIDsInOrder, then authors as before.Verified
EXPLAINuses theembeddingvector key, and the top 3 are identical to the old query for sampled slugs. Under 10ms.go test -p 1 ./internal/...passes against MariaDB 11.8.🤖 Generated with Claude Code
https://claude.ai/code/session_01Q39NMqRSaA8bY8FTaB6Uk5