Ask a folder of markdown notes a question and get an answer where every claim carries the id of
the note it came from, or a plain NOT_IN_NOTES when the notes do not know. It runs on a laptop:
BM25 plus bge-m3 embeddings through a local Ollama, a 3B model for the answer, one SQLite file.
Nothing leaves the machine.
Measured on the demo vault in this repository (22 notes, 98 observations)
with 28 questions: 19 that the notes answer, 6 traps about the same studio that the notes do not
answer, and 3 off-topic traps. Full per-question output: eval/results-demo.json.
| Right note in the top 5 | hybrid 19/19 · vectors 19/19 · BM25 19/19 |
| Answered, with a citation to the expected note | 19 of 19 |
| Citations the model invented | 0 |
| Traps refused | 8 of 9: 3 by the gate before any model call, 5 by the model |
| Real questions refused by mistake | 0 |
| Answer time, ministral-3:3b on an 8 GB M1 | 13 s median |
The miss. Asked "what is Jonas's salary", the model answered "€225" and cited a real note: the monthly cost of a Figma licence that Jonas owns, with the licence renewal date thrown in. The citation check does not catch this, because the id is genuine; the model misread a number that means something else. In an earlier run of this set a stricter refusal rule in the prompt did not fix it and cost one correct answer, so it was reverted. A valid citation is not a correct answer, and on a 3B model this is the failure left.
Hybrid vs vectors on this set. Vectors alone ranked slightly better here (MRR 0.974 against 0.947 for hybrid and 0.776 for BM25 alone). The demo is small and clean; the lexical half is there for vaults full of identifiers such as ticket numbers, product names and codes, where embeddings blur exact strings.
On a real vault. The same retrieval, in an earlier form, runs every day on the author's own notes, which stay private: 517 files, 5,677 observations, a 33 MB index, hybrid search in 78 ms median. On 8 GB of RAM Ollama swaps the embedding model out while the answer model runs, so the first search after an answer can take about five seconds.
- An evidence gate before the model. If the best vector similarity is below the gate, the notes do not cover the question and no model is called. The gate is set from data: off-topic questions scored 0.32–0.34 and the weakest real question 0.50, so the gate sits at 0.42. Traps about the same subject scored 0.49–0.66, overlapping real questions, which is why the gate cannot catch them and the second guard exists.
- Cite or refuse. The prompt requires an observation id after every factual sentence and a fixed refusal token when the sources do not contain the answer.
- A citation check after the model. Every id in the answer is matched against the context that was actually sent. Invented ids are reported, and an answer with no valid citation that is not a refusal is flagged as ungrounded.
- Chunks follow the author's structure, not a token count: a list item with its wrapped lines,
a paragraph, a table row. Each keeps its note title, section and date, including dates from
## Update 2026-09-17headings, so a citation points at one claim. - Ids are a hash of path and content, so re-indexing does not reshuffle them, and stored vectors are carried over instead of recomputed.
- BM25 with a light English and Russian stemmer and pseudo-relevance feedback: the query is re-run with words from the best hits, for when you remember the idea and not the wording. bge-m3 is multilingual, so a question in one language finds a note written in another.
- bge-m3 vectors of the note title, section and text, with a similarity floor so the vector half does not return its "top 60 of anything".
- Reciprocal rank fusion adds the two rankings without putting their scores on one scale.
ollama pull bge-m3
ollama pull ministral-3:3b # or any chat model: --model llama3.2:3b
git clone https://github.com/karusrus/vault-rag && cd vault-rag
python3 -m vaultrag index demo/vault
python3 -m vaultrag embed
python3 -m vaultrag ask "Why did we stop working with Frameline?"
python3 -m vaultrag search "INC-0417"
python3 -m vaultrag eval eval/questions.jsonl --answersPoint index at your own vault. Folders named .obsidian, .git, .trash, _private and
_sensitiveData are never read; add more with --exclude. Python 3.9+, no dependencies beyond
the standard library; pip install numpy makes vector search faster.
- A 3B model misreads numbers, as above. A larger local model is one argument away
(
--model); an API model means replacing one call invaultrag/ask.py. The evaluation above is for ministral-3:3b and nothing else. - The stemmer is crude on purpose. It is good enough for ranking because the vector half covers what it misses, not because it is good.
- The question set is 28 questions written by the author against a fictional vault. It shows the guards work; it is not a benchmark.
- The demo vault is invented: Kestrel Studio, its clients and its people do not exist.
Built out of the author's own second brain, a local search over five thousand notes that an assistant must quote before it may claim anything about the past. The generation half and the evaluation were added so that "RAG" in a profile is something you can open and run.
Ruslan Karymov · AI Enablement & Automation Lead · Creative, marketing and business operations · karusrus.github.io
MIT licence.
