Skip to content
UnknnownnnPublic

About

Yet another fact-knowledge layer designed to make analysis simple and easy

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Fact Knowledge Layer

An evidence-first knowledge layer that converts PDF documents into structured facts, verifies those facts against source text, compares facts across documents, and builds connected knowledge graphs.

Screenshot 2026-09-08 195824

Setup and Run

Prerequisites

  • Python 3.11-3.13. Kuzu currently does not provide a compatible wheel for Python 3.14.
  • Node.js 18 or newer
  • An OpenAI-compatible LLM API key
  • Git

Configure the LLM

Copy the environment template:

Windows:

Copy-Item .env.example .env

Linux:

cp .env.example .env

OpenRouter

LLM_PROVIDER=custom
LLM_API_KEY=your-openrouter-key
LLM_ENDPOINT=https://openrouter.ai/api/v1
LLM_MODEL=your-provider/model-name

Google AI Studio

LLM_PROVIDER=custom
LLM_API_KEY=your-google-ai-studio-key
LLM_ENDPOINT=https://generativelanguage.googleapis.com/v1beta/openai/
LLM_MODEL=gemini-2.5-flash

Cognee is optional and disabled by default. Kuzu remains the native graph database.

Install the Backend

Windows PowerShell:

python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Linux:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Install the Frontend

Windows PowerShell:

Set-Location frontend
npm install
Set-Location ..

Linux:

cd frontend
npm install
cd ..

Run Locally

Start the backend from the repository root:

Windows PowerShell or Linux:

python -m backend.main

In a second terminal, start the frontend:

Windows PowerShell:

Set-Location frontend
npm run dev

Linux:

cd frontend
npm run dev

Open http://localhost:5173.

The backend runs on http://localhost:8000.

Run with Docker

Windows PowerShell or Linux:

docker compose up --build

The frontend is available at http://localhost:3000 and the backend at http://localhost:8000.

What the System Does

PDF Upload
  -> Page-Aware Text Extraction
  -> Chunking
  -> Candidate Fact Discovery
  -> LLM Fact Extraction
  -> Evidence Verification
  -> Fact Normalization
  -> SQLite Storage
  -> Kuzu Graph Synchronization
  -> Fact Comparison
  -> Relationship Classification
  -> UI and Knowledge Graph

The system identifies atomic claims such as financial metrics, dates, percentages, units, reporting periods, scopes, and statuses. Each fact is linked to evidence from the original document.

Approach

Architecture

The application is a single-process FastAPI backend with a React and TypeScript frontend.

  • Frontend: Uploads PDFs and displays documents, facts, evidence, relationships, company trees, and graph views.
  • Backend API: Handles uploads, job status, queries, exports, graph queries, and agent interactions.
  • Extraction pipeline: Parses PDFs, chunks text, discovers candidate chunks, calls the LLM, validates evidence, and normalizes facts.
  • SQLite SQL database: The relational source of truth for documents, pages, facts, evidence, jobs, batches, and relationships.
  • Kuzu graph database: A graph projection used for connected entities, evidence links, relationship traversal, and graph queries.
  • Reasoning layer: Finds related fact pairs and classifies relationships.

Fact Extraction

The PDF is parsed page by page with PyMuPDF. Text is split into page-aware chunks. A lightweight signal detector looks for numbers, percentages, currencies, dates, fiscal periods, and metric terms before sending a chunk to the LLM.

The LLM returns structured JSON containing:

  • Subject and subject type
  • Predicate
  • Value and value type
  • Unit
  • Period and scope
  • Status and approximation flag
  • Confidence
  • Evidence quote

The LLM is used for language interpretation, not as the database or final authority.

Verification and Normalization

The evidence quote is checked against the original chunk using exact, case-insensitive, and fuzzy matching. Unmatched evidence lowers the fact confidence.

Values are normalized before comparison. Examples include converting $2.4 billion to a numeric value, normalizing percentages and units, and converting FY25 to FY2025.

Screenshot 2026-09-08 195909

Reasoning

The reasoning layer uses a staged approach:

  1. Entity compatibility filters unrelated facts.
  2. Embedding similarity finds likely related facts.
  3. Keyword similarity is used when embeddings are unavailable.
  4. Deterministic rules handle obvious matches, conflicts, period differences, scope differences, and unit conversions.
  5. The LLM adjudicates only ambiguous pairs.

Relationships are classified as CORROBORATES, CONTRADICTS, RECONCILES, or UNCERTAIN.

Important Decisions and Trade-offs

  • SQLite plus Kuzu: SQLite provides reliable relational storage and job tracking; Kuzu provides graph traversal. SQLite is authoritative and Kuzu is synchronized from it.
  • Evidence-first extraction: Facts are not treated as trustworthy unless their evidence can be connected to the source text.
  • Deterministic rules before LLM reasoning: This reduces latency, cost, and unnecessary model calls.
  • Incremental batches: Ten-page batches preserve progress and make retries practical when a model request fails.
  • Low concurrency for free models: Serialized extraction is slower but reduces rate limits, timeouts, and malformed responses.
  • Optional Cognee: Cognee is disabled by default because its separate LiteLLM configuration can introduce an unrelated model path. The core Kuzu and SQLite pipeline does not depend on it.

Why These Technologies

Why use both SQLite and Kuzu?

The system has two different data problems:

  1. Reliable record management: Documents, pages, facts, evidence, jobs, and relationships need durable writes, filtering, pagination, foreign-key relationships, and predictable updates.
  2. Connected knowledge exploration: Users need to traverse from a company to its facts, from a fact to its evidence, and from one fact to related or contradictory facts.

SQLite is the right source of truth for the first problem. Kuzu is the right query model for the second. Keeping SQLite authoritative prevents graph traversal concerns from complicating transactional application data, while Kuzu provides a natural representation of the knowledge layer.

This is a deliberate polyglot persistence design: each database is used for the workload it handles best, and synchronization creates a recoverable boundary between them.

Why Kuzu instead of only using relational joins?

The relationships in this application are not limited to fixed reports. A fact can be connected to an entity, a document, evidence, and multiple relationship types such as corroboration, contradiction, reconciliation, or extension.

A relational schema can represent these links, but graph traversal expresses questions like “show all facts connected to this company and the evidence behind them” more directly. Kuzu also provides Cypher-style queries for graph-oriented features while remaining embedded and local to the application.

Why include Cognee?

Cognee is an optional cognitive-memory layer for use cases that go beyond direct fact and relationship queries. When enabled, it can provide:

  • Semantic recall over stored fact statements
  • Dataset-level memory across documents
  • Entity and concept synthesis
  • Multi-hop search and retrieval-oriented workflows

Cognee is intentionally not the primary fact store. The application already has an evidence-aware extraction pipeline and a native Kuzu graph. Cognee is therefore treated as an optional higher-level memory and retrieval service rather than a replacement for SQLite or Kuzu.

Why is Cognee disabled by default?

Cognee internally uses its own LLM and embedding configuration through LiteLLM. That creates a second provider boundary which can silently select a different model, endpoint, or API key than the application-level LLM configuration.

For an evidence-first system, hidden provider configuration is an operational risk. The default path keeps extraction, verification, SQLite persistence, Kuzu synchronization, and reasoning under one explicit configuration. Cognee can be enabled later when its model and endpoint settings are deliberately configured and tested.

Why use an LLM at all?

Rules are effective at detecting patterns such as numbers, percentages, dates, and currencies. They are not sufficient for understanding which entity a value belongs to, what metric is being described, or how a statement is qualified by period and scope.

The LLM handles language interpretation and converts unstructured text into a structured claim. It does not write directly to the graph and it does not determine whether its own evidence is valid. This limits the LLM to the task where it adds the most value.

Why use deterministic reasoning before another LLM call?

Many comparisons are mechanical. Equivalent values, unit conversions, matching periods, and obvious conflicts can be handled more quickly and consistently with rules.

The LLM is reserved for ambiguous comparisons where wording or context requires interpretation. This hybrid strategy reduces cost and latency, makes results easier to audit, and avoids using a probabilistic model for decisions that can be explained mathematically.

Why verify evidence instead of trusting model confidence?

Model confidence is an opinion generated by the model. Evidence verification is an observable check against the source text.

The system therefore stores the claim and the supporting quote together, records the page reference, and reduces confidence when the quote cannot be matched. This does not prove that the source document is true, but it does make the extraction traceable and exposes unsupported model output for review.

Limitations and Next Steps

Current Limitations

  • Scanned or image-only PDFs require OCR; the current parser primarily handles embedded PDF text.
  • LLM output quality depends on the selected provider, model, rate limits, and context window.
  • Free models can be slow or return malformed JSON despite repair and retry handling.
  • Fuzzy evidence matching can preserve weakly supported facts and should not be treated as independent source verification.

Next Steps

  • Add OCR support for scanned PDFs.
  • Add stronger JSON-schema validation and provider-specific adapters.
  • Add automated graph consistency checks between SQLite and Kuzu.
  • Add background queues for multi-user workloads.
  • Add a configurable model fallback chain for provider outages.

Additional Notes

  • SQLite is the source of truth. Kuzu should be treated as a queryable graph projection.
  • Cognee is optional and currently disabled with COGNEE_ENABLED=false.
  • The backend test suite covers parsing, chunking, normalization, evidence validation, LLM response parsing, reasoning, and graph integration.

License

This project is licensed under the MIT License. See LICENSE.

About

Yet another fact-knowledge layer designed to make analysis simple and easy

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages