A modular, observable RAG system for private documents — ingestion, hybrid retrieval, grounded generation with citations, evaluation, and experiment comparison.
DocIntel ingests PDF, TXT, and Markdown files, indexes them for semantic and lexical search, retrieves relevant chunks, generates grounded answers with numbered citations, and records request-level telemetry for every query.
It supports OpenAI and Gemini providers through configuration (LLM_PROVIDER, EMBEDDING_PROVIDER). The same API exposes retrieval-only search, full RAG query, offline evaluation, and side-by-side experiment comparison.
Most RAG demos stop at upload-and-ask. DocIntel adds engineering depth:
- Grounded generation with explicit no-information behavior and citation extraction (including grouped citations like
[Source 1, Source 2]) - Hybrid retrieval — semantic (ChromaDB) + BM25 + RRF, with optional cross-encoder reranking
- Offline evaluation on a fixed 25-question dataset (Recall@K, MRR, LLM-as-judge faithfulness/relevance)
- Experiment comparison between two RAG configurations with persisted metric deltas
- Request observability — traces, aggregate metrics, and a developer dashboard
- Configurable providers — OpenAI or Gemini for embeddings and generation
- Containerized deployment with Docker Compose and health checks
| Area | What is implemented |
|---|---|
| Ingestion | PDF/TXT/MD upload, validation, parsing, deterministic chunking, embeddings, ChromaDB + BM25 indexing |
| Retrieval | Semantic search, BM25, RRF fusion, optional cross-encoder reranking, document filtering |
| Generation | Grounded LLM answers, citation extraction, safe fallback when context is insufficient |
| Evaluation | Fixed 25-question dataset, Recall@K, MRR, LLM-as-judge faithfulness/relevance |
| Experiments | Side-by-side config A vs B comparison with persisted deltas |
| Observability | Request IDs, stage latency, tokens, estimated cost, trace APIs, developer dashboard |
| Engineering | FastAPI modular monolith, SQLite, Docker, structured JSON logging |
flowchart TB
Client[Client / Developer]
FastAPI[FastAPI Application]
Client --> FastAPI
FastAPI --> DocAPI[Document API]
FastAPI --> QueryAPI[Query / Search API]
FastAPI --> TelemetryAPI[Telemetry API]
FastAPI --> EvalAPI[Evaluation API]
FastAPI --> ExpAPI[Experiment API]
FastAPI --> Dash[Developer Dashboard]
DocAPI --> Ingestion[Ingestion Service]
QueryAPI --> RAG[RAG Pipeline]
Ingestion --> Parsers[PDF / TXT / MD Parsers]
Ingestion --> Chunking[Chunking Service]
Ingestion --> Embed[Embedding Provider]
Ingestion --> Chroma[(ChromaDB)]
Ingestion --> BM25[(BM25 Index)]
Ingestion --> SQLite[(SQLite Metadata)]
RAG --> Retrieval[Retrieval Service]
Retrieval --> Chroma
Retrieval --> BM25
Retrieval --> RRF[RRF Fusion]
Retrieval --> Rerank[Cross-Encoder Reranker]
RAG --> LLM[LLM Provider]
RAG --> Citations[Citation Extractor]
RAG --> Traces[(Request Traces)]
EvalAPI --> EvalSvc[Evaluation Service]
EvalSvc --> Retrieval
EvalSvc --> LLM
EvalSvc --> EvalStore[(Evaluation Runs)]
ExpAPI --> ExpSvc[Experiment Service]
ExpSvc --> EvalSvc
ExpSvc --> ExpStore[(Experiment Runs)]
Dash --> SQLite
Dash --> Traces
Dash --> EvalStore
Dash --> ExpStore
flowchart LR
subgraph ingest [Document ingestion]
D[Document] --> Parse[Parse]
Parse --> Chunk[Chunk]
Chunk --> Emb[Embed]
Emb --> VS[(ChromaDB)]
Chunk --> BM[(BM25)]
end
subgraph query [Query path]
Q[Question] --> QEmb[Embed query]
QEmb --> Sem[Semantic search]
Q --> Lex[BM25 search]
Sem --> RRF[RRF optional]
Lex --> RRF
RRF --> Rerank[Rerank optional]
Rerank --> Ctx[Context builder]
Ctx --> Gen[LLM]
Gen --> Ans[Answer + citations]
end
Hybrid retrieval (use_hybrid) and reranking (use_reranking) are optional per request.
| Mode | Flags | Flow |
|---|---|---|
| Semantic only | default | Embed → ChromaDB → threshold → top-K |
| Hybrid | use_hybrid=true |
Semantic + BM25 → RRF → top-K |
| Semantic + rerank | use_reranking=true |
Semantic → rerank pool → cross-encoder → top-K |
| Hybrid + rerank | both true |
Semantic + BM25 → RRF → cross-encoder → top-K |
OpenAI and Gemini embeddings use separate Chroma collections (docintel / docintel_gemini) so vector spaces are never mixed. Switching embedding providers requires re-ingesting documents.
See docs/retrieval.md for score semantics and configuration.
Each /query request persists a trace with retrieval chunks, similarity scores, stage latencies, token counts, estimated cost, citations, and the final answer.
request_id → retrieval → scores → LLM → citations → answer
| Endpoint | Purpose |
|---|---|
GET /api/v1/requests |
Recent request summaries |
GET /api/v1/requests/{id} |
Full request detail |
GET /api/v1/metrics |
Aggregate metrics |
GET /dashboard |
Developer dashboard |
Structured JSON logs are emitted via structlog.
DocIntel includes an offline evaluation engine for the fixed 25-question dataset in evaluation/datasets/sample_eval.json.
Metrics: Recall@K, MRR, faithfulness (LLM-as-judge), answer relevance (LLM-as-judge).
POST /api/v1/evaluations/run
GET /api/v1/evaluations/{id}The evaluation framework supports OpenAI and Gemini when configured. Judge routing follows LLM_PROVIDER (Gemini judge uses the configured Gemini model).
See docs/evaluation_methodology.md.
A real end-to-end evaluation was completed on the 25-question dataset using the Gemini provider:
| Setting | Value |
|---|---|
| Embedding provider | Gemini (gemini-embedding-2, 1536-dim) |
| LLM | gemini-3.6-flash |
| Retrieval | Semantic-only |
| Reranking | Disabled |
| Questions | 25 |
Measured results:
| Metric | Value |
|---|---|
| Recall@5 | 0.3473 |
| MRR | 0.3225 |
| Avg faithfulness | 0.08 |
| Avg answer relevance | 0.076 |
| Avg latency | 2162.90 ms |
| Application estimated cost | $0.00 |
Application-level estimated cost was reported as $0 because Gemini pricing is not configured in the project's pricing table. This is not proof that the Gemini API was free or that provider billing was zero.
The following were successfully validated end-to-end with Gemini configured:
- Document ingestion (PDF/TXT/MD)
- Gemini embeddings → ChromaDB indexing
- ChromaDB semantic retrieval
- Gemini 3.6 Flash generation
- Source citation extraction (single and grouped)
- Request telemetry and trace persistence
- 25-question offline evaluation
Phase 6 controlled experiment — not a fully real OpenAI/Gemini benchmark.
This comparison used:
- Real cross-encoder reranker (
cross-encoder/ms-marco-MiniLM-L-6-v2) - Mocked embeddings
- Mocked LLM generation
- Mocked LLM judge
| Config | Recall@5 | MRR | Avg latency |
|---|---|---|---|
| Semantic only | 1.00 | 0.6400 | 10.53 ms |
| Hybrid + real cross-encoder | 0.80 | 0.6847 | 646.11 ms |
Interpretation: MRR improved slightly (0.64 → 0.68), Recall@5 decreased (1.00 → 0.80), and latency increased substantially (~61×). Hybrid + reranking is a quality-vs-latency tradeoff, not a universal win.
Automated tests use mocked providers for reproducibility. A full real OpenAI evaluation (embeddings + generation + LLM-as-judge) was not performed when no OPENAI_API_KEY was configured.
GET /dashboard — a lightweight Jinja2 dashboard showing:
- Summary cards (requests, success rate, latency, cost)
- Recent request traces
- Latest evaluation run
- Latest experiment comparison
The dashboard displays stored data only; it does not auto-run evaluations or experiments.
See docs/dashboard.md.
| Preview | Description |
|---|---|
| Above (dashboard) | Request activity, evaluations, experiments |
| Above (evaluation) | Real Gemini 25-question evaluation results |
Full gallery (8 screenshots): docs/screenshots.md
| Layer | Technology |
|---|---|
| API | FastAPI, Pydantic, Uvicorn |
| Persistence | SQLite, ChromaDB, on-disk BM25 |
| Embeddings | OpenAI SDK or Gemini API |
| LLM | LiteLLM (OpenAI / Gemini) |
| Retrieval | rank-bm25, sentence-transformers cross-encoder |
| Dashboard | Jinja2, vanilla JS/CSS |
| Testing | pytest, pytest-cov, ruff |
| Container | Docker, Docker Compose, CPU-only PyTorch |
No LangChain, LlamaIndex, or external observability platforms.
DocIntel/
├── app/
│ ├── api/ # FastAPI routes (documents, query, telemetry, eval, experiments, dashboard)
│ ├── db/ # SQLAlchemy engine and repositories
│ ├── models/ # Domain models, Pydantic schemas, config types
│ ├── providers/ # Embeddings, LLM, ChromaDB, BM25, reranker
│ ├── services/ # Ingestion, retrieval, RAG, evaluation, experiments
│ ├── templates/ # Dashboard Jinja2 template
│ └── static/ # Dashboard CSS/JS
├── evaluation/ # Evaluation dataset and sample documents
├── tests/ # 273 automated tests
├── docs/ # Architecture, API, evaluation, design decisions, screenshots
├── Dockerfile
├── docker-compose.yml
└── pyproject.toml
Requires Python 3.11+.
git clone <repo-url>
cd DocIntel
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # macOS/Linux
pip install -e ".[dev]"
cp .env.example .env
# Set OPENAI_API_KEY or GEMINI_API_KEY and provider settings — see .env.example
uvicorn app.main:app --reloadOpen http://127.0.0.1:8000/docs for interactive API docs.
Health check (no API key required):
curl http://127.0.0.1:8000/api/v1/healthcurl -X POST http://127.0.0.1:8000/api/v1/documents \
-F "file=@evaluation/sample_docs/product_guide.txt"
curl -X POST http://127.0.0.1:8000/api/v1/query \
-H "Content-Type: application/json" \
-d '{"question": "What is the refund policy?", "top_k": 5}'cp .env.example .env
docker compose build
docker compose up -d| Property | Value |
|---|---|
| Base image | Python 3.11 slim |
| PyTorch | CPU-only (for cross-encoder reranker) |
| Internal port | 8000 |
| Default host port | 8001 (HOST_PORT override supported) |
| Healthcheck | GET /api/v1/health |
| Volumes | data/chroma, data/uploads, data/bm25, data/sqlite |
Observed on the development environment:
- Image size: 9.21 GB → 2.66 GB (~71% smaller after CPU-only PyTorch)
- Build time: ~18 min → ~7 min
# Override host port when 8001 is occupied
HOST_PORT=18080 docker compose up -dDashboard and API are available at http://localhost:${HOST_PORT:-8001}.
docker compose downSee docs/design_decisions.md for Docker architecture details.
| Endpoint | Purpose |
|---|---|
POST /api/v1/documents |
Upload and ingest a document |
POST /api/v1/search |
Retrieval only (no LLM) |
POST /api/v1/query |
Full RAG with citations |
POST /api/v1/evaluations/run |
Run offline evaluation |
POST /api/v1/experiments/run |
Compare two RAG configs |
GET /api/v1/requests |
List request traces |
GET /api/v1/metrics |
Aggregate telemetry |
GET /dashboard |
Developer dashboard |
GET /api/v1/health |
Health check |
Full route listing: docs/api_reference.md.
| Decision | Rationale |
|---|---|
| Modular monolith | Simple deployment, clear boundaries, appropriate for portfolio scale |
| ChromaDB + BM25 | Local, persistent, no external services |
| SQLite | Simple telemetry/evaluation persistence for single-user workload |
| No LangChain | Direct control over RAG primitives |
| Hybrid retrieval | Semantic captures meaning; BM25 catches exact keywords |
| Cross-encoder reranking | Better ordering on a small candidate pool at significant latency cost |
| CPU-only Docker PyTorch | Avoids ~6 GB of CUDA packages in the container |
| Provider abstraction | OpenAI and Gemini selectable via configuration |
- Single-node local ChromaDB — not a distributed vector database
- SQLite — single-user telemetry and evaluation storage
- Synchronous ingestion and evaluation — no background job queue
- CPU cross-encoder reranking adds latency (~600 ms+ in Phase 6 experiment)
- LLM-as-judge scores have variance; real Gemini faithfulness/relevance on the sample dataset were low (0.08 / 0.076)
- Gemini pricing not in the cost estimator — application may report $0 estimated cost
- Text-only documents (PDF/TXT/MD) — no OCR, audio, or images
- No multi-tenant isolation, streaming, Kubernetes, or cloud deployment
These are conscious scope decisions, not oversights.
- pgvector or Qdrant for production vector storage
- Asynchronous ingestion and evaluation workers
- Larger and domain-specific evaluation datasets
- Gemini/OpenAI pricing in the cost estimator
- Streaming responses
- Stronger citation correctness evaluation
- Multi-tenant document isolation
pytest -q
ruff check .273 tests passing, Ruff clean. 96% line coverage on app/ was measured during Phase 8 validation; re-run pytest --cov=app after major changes to confirm current coverage.
| Document | Description |
|---|---|
| docs/api_reference.md | API routes and examples |
| docs/retrieval.md | Retrieval modes and score semantics |
| docs/evaluation_methodology.md | Metrics, dataset, mocked vs real results |
| docs/design_decisions.md | Architectural choices and tradeoffs |
| docs/experiments.md | Experiment comparison API |
| docs/dashboard.md | Developer dashboard usage |
| docs/screenshots.md | Screenshot gallery |
MIT — see LICENSE. Copyright (c) 2026 Shubham Thakur.

