Skip to content
agastya793Public

About

Production-style RAG knowledge assistant with hybrid retrieval, reranking, evaluation, observability, source citations, experiments, and Dockerized deployment.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

DocIntel

A modular, observable RAG system for private documents — ingestion, hybrid retrieval, grounded generation with citations, evaluation, and experiment comparison.

DocIntel Developer Dashboard

Overview

DocIntel ingests PDF, TXT, and Markdown files, indexes them for semantic and lexical search, retrieves relevant chunks, generates grounded answers with numbered citations, and records request-level telemetry for every query.

It supports OpenAI and Gemini providers through configuration (LLM_PROVIDER, EMBEDDING_PROVIDER). The same API exposes retrieval-only search, full RAG query, offline evaluation, and side-by-side experiment comparison.

Why This Is More Than a PDF Chatbot

Most RAG demos stop at upload-and-ask. DocIntel adds engineering depth:

  • Grounded generation with explicit no-information behavior and citation extraction (including grouped citations like [Source 1, Source 2])
  • Hybrid retrieval — semantic (ChromaDB) + BM25 + RRF, with optional cross-encoder reranking
  • Offline evaluation on a fixed 25-question dataset (Recall@K, MRR, LLM-as-judge faithfulness/relevance)
  • Experiment comparison between two RAG configurations with persisted metric deltas
  • Request observability — traces, aggregate metrics, and a developer dashboard
  • Configurable providers — OpenAI or Gemini for embeddings and generation
  • Containerized deployment with Docker Compose and health checks

Key Features

Area What is implemented
Ingestion PDF/TXT/MD upload, validation, parsing, deterministic chunking, embeddings, ChromaDB + BM25 indexing
Retrieval Semantic search, BM25, RRF fusion, optional cross-encoder reranking, document filtering
Generation Grounded LLM answers, citation extraction, safe fallback when context is insufficient
Evaluation Fixed 25-question dataset, Recall@K, MRR, LLM-as-judge faithfulness/relevance
Experiments Side-by-side config A vs B comparison with persisted deltas
Observability Request IDs, stage latency, tokens, estimated cost, trace APIs, developer dashboard
Engineering FastAPI modular monolith, SQLite, Docker, structured JSON logging

Architecture

flowchart TB
    Client[Client / Developer]
    FastAPI[FastAPI Application]

    Client --> FastAPI

    FastAPI --> DocAPI[Document API]
    FastAPI --> QueryAPI[Query / Search API]
    FastAPI --> TelemetryAPI[Telemetry API]
    FastAPI --> EvalAPI[Evaluation API]
    FastAPI --> ExpAPI[Experiment API]
    FastAPI --> Dash[Developer Dashboard]

    DocAPI --> Ingestion[Ingestion Service]
    QueryAPI --> RAG[RAG Pipeline]

    Ingestion --> Parsers[PDF / TXT / MD Parsers]
    Ingestion --> Chunking[Chunking Service]
    Ingestion --> Embed[Embedding Provider]
    Ingestion --> Chroma[(ChromaDB)]
    Ingestion --> BM25[(BM25 Index)]
    Ingestion --> SQLite[(SQLite Metadata)]

    RAG --> Retrieval[Retrieval Service]
    Retrieval --> Chroma
    Retrieval --> BM25
    Retrieval --> RRF[RRF Fusion]
    Retrieval --> Rerank[Cross-Encoder Reranker]
    RAG --> LLM[LLM Provider]
    RAG --> Citations[Citation Extractor]
    RAG --> Traces[(Request Traces)]

    EvalAPI --> EvalSvc[Evaluation Service]
    EvalSvc --> Retrieval
    EvalSvc --> LLM
    EvalSvc --> EvalStore[(Evaluation Runs)]

    ExpAPI --> ExpSvc[Experiment Service]
    ExpSvc --> EvalSvc
    ExpSvc --> ExpStore[(Experiment Runs)]

    Dash --> SQLite
    Dash --> Traces
    Dash --> EvalStore
    Dash --> ExpStore
Loading

RAG Pipeline

flowchart LR
    subgraph ingest [Document ingestion]
        D[Document] --> Parse[Parse]
        Parse --> Chunk[Chunk]
        Chunk --> Emb[Embed]
        Emb --> VS[(ChromaDB)]
        Chunk --> BM[(BM25)]
    end

    subgraph query [Query path]
        Q[Question] --> QEmb[Embed query]
        QEmb --> Sem[Semantic search]
        Q --> Lex[BM25 search]
        Sem --> RRF[RRF optional]
        Lex --> RRF
        RRF --> Rerank[Rerank optional]
        Rerank --> Ctx[Context builder]
        Ctx --> Gen[LLM]
        Gen --> Ans[Answer + citations]
    end
Loading

Hybrid retrieval (use_hybrid) and reranking (use_reranking) are optional per request.

Retrieval

Mode Flags Flow
Semantic only default Embed → ChromaDB → threshold → top-K
Hybrid use_hybrid=true Semantic + BM25 → RRF → top-K
Semantic + rerank use_reranking=true Semantic → rerank pool → cross-encoder → top-K
Hybrid + rerank both true Semantic + BM25 → RRF → cross-encoder → top-K

OpenAI and Gemini embeddings use separate Chroma collections (docintel / docintel_gemini) so vector spaces are never mixed. Switching embedding providers requires re-ingesting documents.

See docs/retrieval.md for score semantics and configuration.

Observability

Each /query request persists a trace with retrieval chunks, similarity scores, stage latencies, token counts, estimated cost, citations, and the final answer.

request_id → retrieval → scores → LLM → citations → answer
Endpoint Purpose
GET /api/v1/requests Recent request summaries
GET /api/v1/requests/{id} Full request detail
GET /api/v1/metrics Aggregate metrics
GET /dashboard Developer dashboard

Structured JSON logs are emitted via structlog.

Evaluation

DocIntel includes an offline evaluation engine for the fixed 25-question dataset in evaluation/datasets/sample_eval.json.

Metrics: Recall@K, MRR, faithfulness (LLM-as-judge), answer relevance (LLM-as-judge).

POST /api/v1/evaluations/run
GET  /api/v1/evaluations/{id}

The evaluation framework supports OpenAI and Gemini when configured. Judge routing follows LLM_PROVIDER (Gemini judge uses the configured Gemini model).

See docs/evaluation_methodology.md.

Real Gemini Evaluation

A real end-to-end evaluation was completed on the 25-question dataset using the Gemini provider:

Setting Value
Embedding provider Gemini (gemini-embedding-2, 1536-dim)
LLM gemini-3.6-flash
Retrieval Semantic-only
Reranking Disabled
Questions 25

Measured results:

Metric Value
Recall@5 0.3473
MRR 0.3225
Avg faithfulness 0.08
Avg answer relevance 0.076
Avg latency 2162.90 ms
Application estimated cost $0.00

Application-level estimated cost was reported as $0 because Gemini pricing is not configured in the project's pricing table. This is not proof that the Gemini API was free or that provider billing was zero.

Real Gemini Evaluation Response

Verified with Gemini provider

The following were successfully validated end-to-end with Gemini configured:

  • Document ingestion (PDF/TXT/MD)
  • Gemini embeddings → ChromaDB indexing
  • ChromaDB semantic retrieval
  • Gemini 3.6 Flash generation
  • Source citation extraction (single and grouped)
  • Request telemetry and trace persistence
  • 25-question offline evaluation

Retrieval Experiment

Phase 6 controlled experiment — not a fully real OpenAI/Gemini benchmark.

This comparison used:

  • Real cross-encoder reranker (cross-encoder/ms-marco-MiniLM-L-6-v2)
  • Mocked embeddings
  • Mocked LLM generation
  • Mocked LLM judge
Config Recall@5 MRR Avg latency
Semantic only 1.00 0.6400 10.53 ms
Hybrid + real cross-encoder 0.80 0.6847 646.11 ms

Interpretation: MRR improved slightly (0.64 → 0.68), Recall@5 decreased (1.00 → 0.80), and latency increased substantially (~61×). Hybrid + reranking is a quality-vs-latency tradeoff, not a universal win.

Automated tests use mocked providers for reproducibility. A full real OpenAI evaluation (embeddings + generation + LLM-as-judge) was not performed when no OPENAI_API_KEY was configured.

Developer Dashboard

GET /dashboard — a lightweight Jinja2 dashboard showing:

  • Summary cards (requests, success rate, latency, cost)
  • Recent request traces
  • Latest evaluation run
  • Latest experiment comparison

The dashboard displays stored data only; it does not auto-run evaluations or experiments.

See docs/dashboard.md.

Screenshots

Preview Description
Above (dashboard) Request activity, evaluations, experiments
Above (evaluation) Real Gemini 25-question evaluation results

Full gallery (8 screenshots): docs/screenshots.md

Tech Stack

Layer Technology
API FastAPI, Pydantic, Uvicorn
Persistence SQLite, ChromaDB, on-disk BM25
Embeddings OpenAI SDK or Gemini API
LLM LiteLLM (OpenAI / Gemini)
Retrieval rank-bm25, sentence-transformers cross-encoder
Dashboard Jinja2, vanilla JS/CSS
Testing pytest, pytest-cov, ruff
Container Docker, Docker Compose, CPU-only PyTorch

No LangChain, LlamaIndex, or external observability platforms.

Project Structure

DocIntel/
├── app/
│   ├── api/           # FastAPI routes (documents, query, telemetry, eval, experiments, dashboard)
│   ├── db/            # SQLAlchemy engine and repositories
│   ├── models/        # Domain models, Pydantic schemas, config types
│   ├── providers/     # Embeddings, LLM, ChromaDB, BM25, reranker
│   ├── services/      # Ingestion, retrieval, RAG, evaluation, experiments
│   ├── templates/     # Dashboard Jinja2 template
│   └── static/        # Dashboard CSS/JS
├── evaluation/        # Evaluation dataset and sample documents
├── tests/             # 273 automated tests
├── docs/              # Architecture, API, evaluation, design decisions, screenshots
├── Dockerfile
├── docker-compose.yml
└── pyproject.toml

Quick Start

Requires Python 3.11+.

git clone <repo-url>
cd DocIntel

python -m venv .venv
.venv\Scripts\activate          # Windows
# source .venv/bin/activate     # macOS/Linux

pip install -e ".[dev]"

cp .env.example .env
# Set OPENAI_API_KEY or GEMINI_API_KEY and provider settings — see .env.example

uvicorn app.main:app --reload

Open http://127.0.0.1:8000/docs for interactive API docs.

Health check (no API key required):

curl http://127.0.0.1:8000/api/v1/health

Example: upload and query

curl -X POST http://127.0.0.1:8000/api/v1/documents \
  -F "file=@evaluation/sample_docs/product_guide.txt"

curl -X POST http://127.0.0.1:8000/api/v1/query \
  -H "Content-Type: application/json" \
  -d '{"question": "What is the refund policy?", "top_k": 5}'

Docker

cp .env.example .env
docker compose build
docker compose up -d
Property Value
Base image Python 3.11 slim
PyTorch CPU-only (for cross-encoder reranker)
Internal port 8000
Default host port 8001 (HOST_PORT override supported)
Healthcheck GET /api/v1/health
Volumes data/chroma, data/uploads, data/bm25, data/sqlite

Observed on the development environment:

  • Image size: 9.21 GB → 2.66 GB (~71% smaller after CPU-only PyTorch)
  • Build time: ~18 min → ~7 min
# Override host port when 8001 is occupied
HOST_PORT=18080 docker compose up -d

Dashboard and API are available at http://localhost:${HOST_PORT:-8001}.

docker compose down

See docs/design_decisions.md for Docker architecture details.

API

Endpoint Purpose
POST /api/v1/documents Upload and ingest a document
POST /api/v1/search Retrieval only (no LLM)
POST /api/v1/query Full RAG with citations
POST /api/v1/evaluations/run Run offline evaluation
POST /api/v1/experiments/run Compare two RAG configs
GET /api/v1/requests List request traces
GET /api/v1/metrics Aggregate telemetry
GET /dashboard Developer dashboard
GET /api/v1/health Health check

Full route listing: docs/api_reference.md.

Design Tradeoffs

Decision Rationale
Modular monolith Simple deployment, clear boundaries, appropriate for portfolio scale
ChromaDB + BM25 Local, persistent, no external services
SQLite Simple telemetry/evaluation persistence for single-user workload
No LangChain Direct control over RAG primitives
Hybrid retrieval Semantic captures meaning; BM25 catches exact keywords
Cross-encoder reranking Better ordering on a small candidate pool at significant latency cost
CPU-only Docker PyTorch Avoids ~6 GB of CUDA packages in the container
Provider abstraction OpenAI and Gemini selectable via configuration

Limitations

  • Single-node local ChromaDB — not a distributed vector database
  • SQLite — single-user telemetry and evaluation storage
  • Synchronous ingestion and evaluation — no background job queue
  • CPU cross-encoder reranking adds latency (~600 ms+ in Phase 6 experiment)
  • LLM-as-judge scores have variance; real Gemini faithfulness/relevance on the sample dataset were low (0.08 / 0.076)
  • Gemini pricing not in the cost estimator — application may report $0 estimated cost
  • Text-only documents (PDF/TXT/MD) — no OCR, audio, or images
  • No multi-tenant isolation, streaming, Kubernetes, or cloud deployment

These are conscious scope decisions, not oversights.

Future Improvements

  • pgvector or Qdrant for production vector storage
  • Asynchronous ingestion and evaluation workers
  • Larger and domain-specific evaluation datasets
  • Gemini/OpenAI pricing in the cost estimator
  • Streaming responses
  • Stronger citation correctness evaluation
  • Multi-tenant document isolation

Testing

pytest -q
ruff check .

273 tests passing, Ruff clean. 96% line coverage on app/ was measured during Phase 8 validation; re-run pytest --cov=app after major changes to confirm current coverage.

Documentation

Document Description
docs/api_reference.md API routes and examples
docs/retrieval.md Retrieval modes and score semantics
docs/evaluation_methodology.md Metrics, dataset, mocked vs real results
docs/design_decisions.md Architectural choices and tradeoffs
docs/experiments.md Experiment comparison API
docs/dashboard.md Developer dashboard usage
docs/screenshots.md Screenshot gallery

License

MIT — see LICENSE. Copyright (c) 2026 Shubham Thakur.

About

Production-style RAG knowledge assistant with hybrid retrieval, reranking, evaluation, observability, source citations, experiments, and Dockerized deployment.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages