Diagnose why your RAG retrieval is returning the wrong documents.
PyVectorHound is a component-level diagnostic engine for retrieval-augmented generation (RAG) pipelines. Point it at a set of search results (from your own pipeline, or from a live Qdrant/Chroma/Milvus/pgvector/Weaviate instance) and it isolates which stage is failing — embedding quality or vector search ranking — and gives you plain-English, ranked recommendations.
- A RAG pipeline is returning irrelevant documents and you don't know
whether it's the embeddings or the ranking —
Diagnosis/Hound.diagnose()isolates which stage is at fault instead of forcing you to eyeball cosine-similarity scores. - Catching retrieved documents that look similar but contradict the query — the LLM-as-judge faithfulness check, when pure distance metrics would score them as a good match.
- Comparing embedding models or database configs objectively — track
metrics over time with
Hound.track_metric()so quality thresholds are relative to your own baseline, not a hardcoded cutoff. - Not yet a good fit for: diagnosing BM25/keyword-search or reranker
stages specifically (both report
UNKNOWN— not implemented, not guessed); a fully automated model-comparison quality score without supplying your ownquality_fn.
PyVectorHound does not run an embedding model, a reranker, or BM25 for you, and it is not a vector database. It's a diagnostic layer that sits on top of retrieval results you already have (or that it fetches from your vector DB) and tells you, with real computed metrics, what's wrong:
- Embedding-space diagnostics (isotropy, coverage, distinctiveness) —
computed by a Rust extension (
pyvectorhound._core, built via PyO3) from the real per-document embeddings your database adapter'sget_embeddings()returns. - Vector search accuracy (precision, recall, MRR) — computed against
expected_docsyou supply as ground truth. - LLM-as-judge faithfulness / contradiction checking — pass
document_textsand anllm_judge_fn(no LLM client bundled, same pattern asembed_fn) andDiagnosiswill flag retrieved documents that score well on embedding similarity but are actually irrelevant to or contradict the query — the failure mode pure distance metrics can't see. - Root cause + ranked recommendations — plain-English output combining the above.
- Concurrent batch evaluation —
Hound.diagnose_batch()runs many queries at once on a thread pool (with automatic retry/backoff onembed_fn/llm_judge_fncalls) instead of one at a time, for large-scale evaluation runs. - Model-agnostic quality thresholds — once you've tracked a few
diagnoses with
Hound.track_metric(), GOOD/MODERATE/WEAK status is computed relative to your own historical baseline instead of a fixed cutoff, so it doesn't drift when you switch embedding models.
BM25 (keyword search) and reranker diagnostics are reported as "UNKNOWN":
PyVectorHound doesn't run a keyword-search index or a reranker itself, and
Diagnosis doesn't yet accept external BM25/reranker scores as input, so
rather than fabricate a number for a component it can't measure, it says so.
If a component doesn't have enough input to measure honestly (no adapter,
fewer than 2 documents with embeddings, or no expected_docs), it's
reported as "UNKNOWN" with an explanation of what to supply — never a
made-up number.
PyVectorHound does not bundle an embedding model. If you want
Hound.diagnose() to embed your query text for you, pass it an embed_fn
(a thin wrapper around whatever you already use — OpenAI, Cohere,
sentence-transformers, etc.). Without one, pass a precomputed
query_embedding per call. It will not silently generate a random vector
and pretend the resulting diagnosis means something.
pyvectorhound.rank_and_diversify(results, top_k) combines BM25, semantic,
and recency signals (whichever you supply per result) into a single
multi-criteria ranking, then diversifies the top results using a real
cosine-similarity penalty computed from each document's embedding — a
near-duplicate of an already-selected result gets pulled down, a result
that's orthogonal to everything selected so far pays no penalty. (This used
to be a hardcoded 10%-per-result placeholder that never looked at the
embeddings at all; it's now backed by src/retrieval_ranking.rs's real
cosine_similarity.) pyvectorhound.compute_reranker_metrics(...) then
gives you MRR/NDCG/Precision@K/Recall@K/diversity against your own
ground-truth relevant-doc list.
from pyvectorhound import rank_and_diversify, compute_reranker_metrics
results = [
("doc_1", {"semantic": 0.9}, [1.0, 0.0, 0.3]), # (doc_id, scores, embedding)
("doc_2", {"semantic": 0.85}, [0.98, 0.1, 0.25]), # near-duplicate of doc_1
("doc_3", {"semantic": 0.6}, [0.0, 1.0, 0.0]),
]
ranked = rank_and_diversify(results, top_k=3)
metrics = compute_reranker_metrics(
[r["document_id"] for r in ranked],
[r["diversity_score"] for r in ranked],
relevant_docs=["doc_1", "doc_3"],
k=3,
)src/quantization.rs adds per-vector int8 scalar quantization for shrinking
the memory footprint of large local vector indices. Each f32 vector is
mapped linearly to the i8 range (-128..=127) using that vector's own
min/max, alongside a (scale, offset) pair needed to dequantize it back to
an approximate f32 vector. This is scalar quantization, not product
quantization, and it isn't SIMD/GPU accelerated — both are out of scope for
this feature.
Tradeoff: storing i8 instead of f32 is a 4x memory reduction
for the vector data (1 byte/dim vs 4 bytes/dim), at the cost of bounded
reconstruction error — max absolute error per element is at most
(max - min) / 255 for the vector being quantized (half a quantization
step in practice). For a typical normalized embedding in [-1, 1], that's
a worst-case error bound of about 0.0078 per dimension; measured max
error on a random 384-dim [-1, 1] vector in the test suite was ~0.0039.
Constant vectors (including all-zero) round-trip exactly.
Exposed to Python via pyvectorhound._core:
from pyvectorhound import _core
# Single vector
result = _core.py_quantize_vector([0.12, -0.87, 0.5, 0.0])
# {"values": [i8, ...], "scale": float, "offset": float, "min": float, "max": float}
restored = _core.py_dequantize_vector(result["values"], result["scale"], result["offset"])
# Batch (one scale/offset per vector)
batch = _core.py_quantize_batch([[0.1, 0.2], [-1.0, 1.0]])
# {"values": [[i8, ...], ...], "params": [{"scale":..., "offset":..., "min":..., "max":...}, ...]}
restored_batch = _core.py_dequantize_batch(
batch["values"],
[p["scale"] for p in batch["params"]],
[p["offset"] for p in batch["params"]],
)This is currently a low-level building block (quantize/dequantize
primitives) rather than integrated into Hound's storage/retrieval path —
if you want quantized on-disk storage for your own index today, call these
functions directly.
Being upfront about what's still a stub, rather than leaving it to look finished:
ModelComparison/Hound.compare_models()reports real, published cost/latency metadata for known models, but has no way to measure quality (F1/NDCG) on its own — passquality_fnfor real numbers, or it reports quality as unmeasured.Hound.compare_metrics()andHound.detect_drift()raiseNotImplementedError;QualityScorer.trend_analysis()instead returns a dict with"direction": "unknown"and an explanation. None of the three has a historical data store to compute a real trend from. UseHound.track_metric()+Hound.get_trend_report()(backed by the real, testedTrendAnalyzer) instead.- As of v1.5.0 (current), PyPI carries a macOS arm64 / CPython 3.11 wheel
plus a source distribution (
sdist). The wheel is still single-platform/single-ABI (noabi3build yet — see below), but the sdist meanspip install pyvectorhoundon any other interpreter or OS now builds the current version from source instead of silently falling back to an old release, as long as a Rust toolchain is available locally. If the sdist build fails for you, install straight from the repo instead:pip install git+https://github.com/Mullassery/PyVectorHound.git. A proper multi-platform wheel matrix (built in CI,abi3so one wheel covers multiple CPython versions) is still open work. - GitHub Actions CI (the badge above) is green as of this pass — a prior
revision of this note described the
dtolnay/rust-toolchainsetup step as broken (missingtoolchaininput), butci.ymlalready uses the@stableform, which doesn't need one; CI has passed on every push since the fix. The full test suite passes (174/174 as of this writing, run viapytest tests/ -vwith the Rust extension built). - OpenTelemetry / LangChain / LlamaIndex / MCP integrations, the CLI, and
the REST server exist and have passing tests but have seen far less
real-world use than the core
Hound/Diagnosispath above.
pip install pyvectorhoundOptional vector database clients (only install the one(s) you use):
pip install pyvectorhound[qdrant] # Qdrant
pip install pyvectorhound[chroma] # Chroma
pip install pyvectorhound[milvus] # Milvus
pip install pyvectorhound[weaviate] # Weaviate
pip install pyvectorhound[pgvector] # PostgreSQL + pgvectorRequires Python 3.8+.
This is the fastest way to try it — no live database or embedding model
needed. Diagnosis fetches per-document embeddings for you via a small
adapter object (anything with a get_embeddings(doc_ids) -> dict method);
without one, the embedding component honestly reports "UNKNOWN" instead
of a fabricated score.
from pyvectorhound import Diagnosis
class InMemoryAdapter:
"""Anything with get_embeddings(doc_ids) works -- swap in your own
QdrantAdapter/ChromaAdapter/etc., or a wrapper around your pipeline."""
def __init__(self, embeddings_by_id):
self._embeddings_by_id = embeddings_by_id
def get_embeddings(self, doc_ids):
return {d: self._embeddings_by_id[d] for d in doc_ids if d in self._embeddings_by_id}
results = [
{"id": "pricing.pdf", "score": 0.91},
{"id": "onboarding.md", "score": 0.84},
{"id": "faq.md", "score": 0.79},
]
diagnosis = Diagnosis(
query="What's your return policy?",
results=results,
expected_docs=["returns.pdf", "policy.md"], # ground truth
adapter=InMemoryAdapter(my_document_embeddings),
)
diagnosis.analyze()
print(diagnosis.root_cause())
for rec in diagnosis.recommendations():
print(f"[{rec['priority']}] {rec['action']}")
print(diagnosis.hunt()) # full plain-English reportA runnable version (with synthetic embeddings so it works with no setup) is
in examples/retrieval_debug.py.
from pyvectorhound import Hound
hound = Hound(
db="qdrant", # qdrant | chroma | milvus | weaviate | postgres
endpoint="localhost:6333",
index_name="documents",
# PyVectorHound doesn't ship an embedding model -- wrap whatever you use:
embed_fn=lambda text: my_embedding_client.embed(text),
)
diagnosis = hound.diagnose(
query="What's your return policy?",
expected_docs=["returns.pdf", "policy.md"],
top_k=5,
)
print(diagnosis.hunt())Hound connects lazily — constructing it doesn't require a live server,
only calling diagnose() (or another querying method) does. diagnose()
already passes self.adapter into Diagnosis, so embedding-space
diagnostics work out of the box against your real database.
| Component | What it measures | Requires |
|---|---|---|
| Embedding | Isotropy, coverage, distinctiveness of the retrieved documents' real embeddings | An adapter with get_embeddings(), and ≥2 retrieved documents |
| Vector search | Precision, recall, MRR | expected_docs (ground truth) |
| Faithfulness | LLM-judge contradiction/relevance check | document_texts + llm_judge_fn |
| BM25 (keyword) | Not implemented — reports UNKNOWN |
n/a |
| Reranker | Not implemented — reports UNKNOWN |
n/a |
Every measured component is computed for real from the input you give it; nothing is guessed when the input isn't there.
from pyvectorhound import Hound
def judge(query: str, doc_texts: list[str]) -> dict:
# Wrap whatever LLM client you already use -- PyVectorHound doesn't
# bundle one. Must return at least a "contradiction_score" (0.0-1.0,
# lower is more faithful) and/or "faithful"/"contradicted_count".
response = my_llm_client.judge_faithfulness(query, doc_texts)
return {
"contradiction_score": response.score,
"contradicted_count": response.contradicted,
"reasoning": response.explanation,
}
hound = Hound(db="qdrant", embed_fn=my_embed_fn, llm_judge_fn=judge)
diagnosis = hound.diagnose(
query="What's your return policy?",
document_texts={"returns.pdf": "...", "policy.md": "..."}, # doc_id -> text
)
print(diagnosis.metrics()["faithfulness"])
# Evaluate many queries concurrently instead of one at a time:
diagnoses = hound.diagnose_batch(
queries=["query 1", "query 2", "query 3"],
document_texts=[{"a": "..."}, None, {"b": "..."}], # per-query, optional
max_workers=8,
)Track a few diagnoses over time and quality-status classification switches from a fixed cutoff to your own historical baseline automatically:
hound.track_metric("vector_search_precision", diagnosis.metrics()["vector_search"]["precision"])
# After ~5+ tracked points, later diagnose() calls classify status
# (GOOD/MODERATE/WEAK) relative to that baseline instead of a fixed number.Ragas is the closest OSS competitor — the standard for RAG evaluation
metrics. We ran both against the same real corpus (168 real chunks from 3
live English Wikipedia articles: Transformer architecture, attention
mechanisms, retrieval-augmented generation), real embeddings
(nomic-embed-text via a local Ollama instance, not fabricated vectors),
and a real local LLM judge/generator (qwen2.5:7b-instruct via Ollama — no
cloud API keys used or needed). 4 real queries, including 2 deliberately
hard cases (one genuinely unanswerable from the corpus, one prone to
retrieving a topically-similar-but-wrong article).
| Query | PyVectorHound vector_search |
PyVectorHound faithfulness |
Ragas context_precision |
Ragas context_recall |
Ragas faithfulness |
|---|---|---|---|---|---|
| "easy" (attention mechanism) | WEAK — 0% precision/recall (retrieved 0/4 chunks from the expected article) | GOOD (retrieved passage still explains the mechanism) | 0.9999 | 1.0 | 1.0 |
| "cross_chunk" (how RAG reduces hallucination) | GOOD — 100% precision/recall | GOOD (retrieved passages are on-topic) | 0.9999 | 1.0 | 0.0 (generated answer refused to use the context) |
| "no_answer" (OpenAI's exact founding date — not in corpus) | UNKNOWN (correctly declines to guess without ground truth) | GOOD (0 faithful/0 contradicted — no false positive) | 0.0 (correctly flags irrelevant context) | 1.0 | 1.0 |
| "distractor" (self- vs cross-attention) | WEAK — 50% precision | GOOD | 0.9999 | 1.0 | 1.0 |
What this actually shows, not just the numbers: on the "easy" and
"distractor" queries, PyVectorHound's vector_search metric scored the
retrieval as weak/failing — but that's a real limitation of this test's
ground truth, not necessarily a real retrieval failure: the expected-docs
ground truth was defined by which Wikipedia article a chunk came from,
and in both cases the retrieved chunks (from a different article) still
genuinely contained the answer, confirmed by both LLM judges. Ragas's
context_precision/context_recall, which score against a written
reference answer rather than document identity, aren't fooled by that —
a real methodological lesson: document-identity ground truth is brittle,
answer-based ground truth is more robust. That's a real tradeoff of
PyVectorHound's approach, not a bug we found and fixed.
The "cross_chunk" row shows the real, useful architectural difference
between the two tools: PyVectorHound's faithfulness check judges the
retrieved passages (are they relevant/non-contradictory?) and correctly
said GOOD — they were. Ragas's faithfulness judges the generated
answer's claims against the context, and correctly caught that the LLM's
actual answer declined to use good context and hedged instead — a real
generation-layer failure that PyVectorHound has no visibility into,
because it never requires an LLM to generate a final answer at all.
That's the real scope difference: PyVectorHound diagnoses the retrieval
layer standalone (useful when you don't have or don't want to run a full
generation step); Ragas evaluates the full retrieve-then-generate
pipeline end to end, which needs a real generated answer and a written
reference answer to run at all.
One real, reproducible gotcha found while building this benchmark:
PyVectorHound's ChromaAdapter.connect() (database.py) calls
chromadb.Client() with no persistence path, so it only works when
populated in the same process — a separate script/process using its own
chromadb.Client() cannot see data another process wrote, since Chroma's
ephemeral in-memory client isn't shared across process boundaries (it is
shared across separate Client() calls within one process). Not a code
bug — this is real, documented Chroma default behavior — but worth
knowing before assuming endpoint="" means "persistent local store."
hound.quality_scorer()—QualityScorerfor scoring an embedding's validity, and (given corpus neighbors via the adapter) real isotropy/coverage/distinctiveness against the corpus.hound.benchmark()—PerformanceBenchmarkfor latency percentiles and database/embedding-model comparisons.hound.analyze_trends()—TrendAnalyzerfor tracking metrics over time and detecting drift, regressions, and anomalies from real tracked values.hound.tracer()/hound.replayer()— capture a retrieval pipeline run and replay it under different configurations to compare recall/latency.
See examples/ for runnable scripts, and
docs/ARCHITECTURE.md / docs/GUIDE.md
for more detail. See ROADMAP_HONEST.md for a plain
list of what's not built, what's untested, and known technical debt.
git clone https://github.com/Mullassery/PyVectorHound.git
cd PyVectorHound
pip install maturin
maturin develop --release # builds the Rust extension in place
pip install -e ".[dev]"
pytest tests/ -vThis project is licensed under the Apache License 2.0.