An evidence-first knowledge layer that converts PDF documents into structured facts, verifies those facts against source text, compares facts across documents, and builds connected knowledge graphs.
- Python 3.11-3.13. Kuzu currently does not provide a compatible wheel for Python 3.14.
- Node.js 18 or newer
- An OpenAI-compatible LLM API key
- Git
Copy the environment template:
Windows:
Copy-Item .env.example .envLinux:
cp .env.example .envLLM_PROVIDER=custom
LLM_API_KEY=your-openrouter-key
LLM_ENDPOINT=https://openrouter.ai/api/v1
LLM_MODEL=your-provider/model-nameLLM_PROVIDER=custom
LLM_API_KEY=your-google-ai-studio-key
LLM_ENDPOINT=https://generativelanguage.googleapis.com/v1beta/openai/
LLM_MODEL=gemini-2.5-flashCognee is optional and disabled by default. Kuzu remains the native graph database.
Windows PowerShell:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txtLinux:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtWindows PowerShell:
Set-Location frontend
npm install
Set-Location ..Linux:
cd frontend
npm install
cd ..Start the backend from the repository root:
Windows PowerShell or Linux:
python -m backend.mainIn a second terminal, start the frontend:
Windows PowerShell:
Set-Location frontend
npm run devLinux:
cd frontend
npm run devOpen http://localhost:5173.
The backend runs on http://localhost:8000.
Windows PowerShell or Linux:
docker compose up --buildThe frontend is available at http://localhost:3000 and the backend at http://localhost:8000.
PDF Upload
-> Page-Aware Text Extraction
-> Chunking
-> Candidate Fact Discovery
-> LLM Fact Extraction
-> Evidence Verification
-> Fact Normalization
-> SQLite Storage
-> Kuzu Graph Synchronization
-> Fact Comparison
-> Relationship Classification
-> UI and Knowledge Graph
The system identifies atomic claims such as financial metrics, dates, percentages, units, reporting periods, scopes, and statuses. Each fact is linked to evidence from the original document.
The application is a single-process FastAPI backend with a React and TypeScript frontend.
- Frontend: Uploads PDFs and displays documents, facts, evidence, relationships, company trees, and graph views.
- Backend API: Handles uploads, job status, queries, exports, graph queries, and agent interactions.
- Extraction pipeline: Parses PDFs, chunks text, discovers candidate chunks, calls the LLM, validates evidence, and normalizes facts.
- SQLite SQL database: The relational source of truth for documents, pages, facts, evidence, jobs, batches, and relationships.
- Kuzu graph database: A graph projection used for connected entities, evidence links, relationship traversal, and graph queries.
- Reasoning layer: Finds related fact pairs and classifies relationships.
The PDF is parsed page by page with PyMuPDF. Text is split into page-aware chunks. A lightweight signal detector looks for numbers, percentages, currencies, dates, fiscal periods, and metric terms before sending a chunk to the LLM.
The LLM returns structured JSON containing:
- Subject and subject type
- Predicate
- Value and value type
- Unit
- Period and scope
- Status and approximation flag
- Confidence
- Evidence quote
The LLM is used for language interpretation, not as the database or final authority.
The evidence quote is checked against the original chunk using exact, case-insensitive, and fuzzy matching. Unmatched evidence lowers the fact confidence.
Values are normalized before comparison. Examples include converting $2.4 billion to a numeric value, normalizing percentages and units, and converting FY25 to FY2025.
The reasoning layer uses a staged approach:
- Entity compatibility filters unrelated facts.
- Embedding similarity finds likely related facts.
- Keyword similarity is used when embeddings are unavailable.
- Deterministic rules handle obvious matches, conflicts, period differences, scope differences, and unit conversions.
- The LLM adjudicates only ambiguous pairs.
Relationships are classified as CORROBORATES, CONTRADICTS, RECONCILES, or UNCERTAIN.
- SQLite plus Kuzu: SQLite provides reliable relational storage and job tracking; Kuzu provides graph traversal. SQLite is authoritative and Kuzu is synchronized from it.
- Evidence-first extraction: Facts are not treated as trustworthy unless their evidence can be connected to the source text.
- Deterministic rules before LLM reasoning: This reduces latency, cost, and unnecessary model calls.
- Incremental batches: Ten-page batches preserve progress and make retries practical when a model request fails.
- Low concurrency for free models: Serialized extraction is slower but reduces rate limits, timeouts, and malformed responses.
- Optional Cognee: Cognee is disabled by default because its separate LiteLLM configuration can introduce an unrelated model path. The core Kuzu and SQLite pipeline does not depend on it.
The system has two different data problems:
- Reliable record management: Documents, pages, facts, evidence, jobs, and relationships need durable writes, filtering, pagination, foreign-key relationships, and predictable updates.
- Connected knowledge exploration: Users need to traverse from a company to its facts, from a fact to its evidence, and from one fact to related or contradictory facts.
SQLite is the right source of truth for the first problem. Kuzu is the right query model for the second. Keeping SQLite authoritative prevents graph traversal concerns from complicating transactional application data, while Kuzu provides a natural representation of the knowledge layer.
This is a deliberate polyglot persistence design: each database is used for the workload it handles best, and synchronization creates a recoverable boundary between them.
The relationships in this application are not limited to fixed reports. A fact can be connected to an entity, a document, evidence, and multiple relationship types such as corroboration, contradiction, reconciliation, or extension.
A relational schema can represent these links, but graph traversal expresses questions like “show all facts connected to this company and the evidence behind them” more directly. Kuzu also provides Cypher-style queries for graph-oriented features while remaining embedded and local to the application.
Cognee is an optional cognitive-memory layer for use cases that go beyond direct fact and relationship queries. When enabled, it can provide:
- Semantic recall over stored fact statements
- Dataset-level memory across documents
- Entity and concept synthesis
- Multi-hop search and retrieval-oriented workflows
Cognee is intentionally not the primary fact store. The application already has an evidence-aware extraction pipeline and a native Kuzu graph. Cognee is therefore treated as an optional higher-level memory and retrieval service rather than a replacement for SQLite or Kuzu.
Cognee internally uses its own LLM and embedding configuration through LiteLLM. That creates a second provider boundary which can silently select a different model, endpoint, or API key than the application-level LLM configuration.
For an evidence-first system, hidden provider configuration is an operational risk. The default path keeps extraction, verification, SQLite persistence, Kuzu synchronization, and reasoning under one explicit configuration. Cognee can be enabled later when its model and endpoint settings are deliberately configured and tested.
Rules are effective at detecting patterns such as numbers, percentages, dates, and currencies. They are not sufficient for understanding which entity a value belongs to, what metric is being described, or how a statement is qualified by period and scope.
The LLM handles language interpretation and converts unstructured text into a structured claim. It does not write directly to the graph and it does not determine whether its own evidence is valid. This limits the LLM to the task where it adds the most value.
Many comparisons are mechanical. Equivalent values, unit conversions, matching periods, and obvious conflicts can be handled more quickly and consistently with rules.
The LLM is reserved for ambiguous comparisons where wording or context requires interpretation. This hybrid strategy reduces cost and latency, makes results easier to audit, and avoids using a probabilistic model for decisions that can be explained mathematically.
Model confidence is an opinion generated by the model. Evidence verification is an observable check against the source text.
The system therefore stores the claim and the supporting quote together, records the page reference, and reduces confidence when the quote cannot be matched. This does not prove that the source document is true, but it does make the extraction traceable and exposes unsupported model output for review.
- Scanned or image-only PDFs require OCR; the current parser primarily handles embedded PDF text.
- LLM output quality depends on the selected provider, model, rate limits, and context window.
- Free models can be slow or return malformed JSON despite repair and retry handling.
- Fuzzy evidence matching can preserve weakly supported facts and should not be treated as independent source verification.
- Add OCR support for scanned PDFs.
- Add stronger JSON-schema validation and provider-specific adapters.
- Add automated graph consistency checks between SQLite and Kuzu.
- Add background queues for multi-user workloads.
- Add a configurable model fallback chain for provider outages.
- SQLite is the source of truth. Kuzu should be treated as a queryable graph projection.
- Cognee is optional and currently disabled with
COGNEE_ENABLED=false. - The backend test suite covers parsing, chunking, normalization, evidence validation, LLM response parsing, reasoning, and graph integration.
This project is licensed under the MIT License. See LICENSE.