DaveLLM is an Electron desktop client backed by a FastAPI router. The router discovers models from configured Ollama nodes, streams chat responses from Ollama's native chat API, and persists conversations, projects, vector indexes, feedback, performance, and cost data on the router host.
| Document | Authority |
|---|---|
| README.md | Primary setup, configuration, feature, and validation guide. |
| PROJECT_SPEC.md | Current product scope, functional requirements, architecture, security, limits, and acceptance criteria. |
| INTEGRATION.md | Authoritative frontend/backend flow, authentication, and API contracts. |
| CLAUDE.md | Maintainer architecture, trust boundaries, and repository constraints. |
| docs/OPERATIONS.md | Placeholder-only launch, health, inventory, persistence, and troubleshooting runbook. |
| docs/DAVELLM_TOOLS.md | Tool settings, the qualified tools, and the extended tools' arguments, limits, path rules, and approved edits. |
| docs/DAVELLM_GIT_SECURITY_REVIEW.md | Independent security review of the read-only Git tools, its findings, and their fixes. |
| docs/DAVEHARNESS_CAPABILITIES.md | Generated capabilities manifest: every tool, flag, run budget, host limit, /tools route, and tool-module constant, with a JSON twin for tooling. |
| DaveHarness boundary and versioning decision | Implemented product ownership, package seam, SemVer authority, and revisit triggers. |
| DaveHarness implementation plan | Authoritative 60-action roadmap through shipped 0.2.0 execution semantics, 0.3.0 leaf contracts, 0.4.0 policy and budgets, 0.5.0 run state, 0.6.0 cancellation/deadlines, 0.7.0 events, 0.8.0 facade, and 0.9.0 DaveLLM integration toward a qualified 1.0.0 contract. |
| dave-llm-feature-analysis-2026-03-25.md | Dated feature-status analysis with explicit verification boundaries. |
| docs/EXECUTABLE_PROMPT_SERIES.md | P0 through P8 decision, plan, implementation, and review contracts. |
| Public landing page | Published product overview and quickstart; not the desktop runtime static root. |
- Python 3
- Node.js and npm
- One or more reachable Ollama nodes
DAVE_API_KEYset to a non-empty local secret
The repository does not include models, Whisper assets, credentials, or runtime data. On macOS, install the optional local dictation runtime after the normal setup:
npm run install:whisper:macosThis installs Homebrew whisper-cpp when needed and downloads the checksum-verified tiny.en model to ~/Library/Application Support/DaveLLM/models/. The model remains app-managed runtime data and is never committed.
python3 -m venv venv
source venv/bin/activate
python -m pip install -r requirements.txt
npm ciConfigure the process environment before starting. Values below are placeholders, not working cluster inventory:
export DAVE_API_KEY='<local-secret>'
export DAVE_NODES='[{"id":"<node-id>","name":"<display-name>","url":"http://<ollama-host>:11434"}]'
npm startElectron starts uvicorn, waits up to 15 seconds for public /health, injects X-API-Key only into requests to its exact loopback backend origin, and then loads the UI. The default Python is the repository virtual environment (venv/bin/python on macOS and Linux, venv\Scripts\python.exe on Windows). Override the Python executable with DAVE_PYTHON, the port with DAVE_PORT, or the readiness timeout with DAVE_STARTUP_TIMEOUT_MS.
Install a Keychain-backed launcher after completing the normal Python and npm setup:
npm run install:macosThe installer creates a strong DAVE_API_KEY in macOS Keychain when the com.davellm.api-key item does not already exist, preserves an existing item, and installs ~/Applications/DaveLLM Launcher.app. The launcher requires the Tailscale peers named dominic and walter, then adds duncan only when that peer is online and its Ollama /api/tags endpoint responds within four seconds. An unavailable Duncan is logged and skipped without blocking startup. The launcher constructs DAVE_NODES in memory, with a capability profile for each node (DL-ROUTE-01, below), uses ~/Library/Application Support/DaveLLM for persistence, and starts the existing Electron application. It also sets DAVE_ENABLE_TOOLS=true and DAVE_ENABLE_EXTENDED_TOOLS=true so the Run tools button works, and sets DAVE_SEARCH_URL when a SearXNG instance on Dominic (port 8890, or DAVE_SEARCH_PORT) answers its health check; export either flag as false before launching to turn it off. It does not enable shell.exec or set DAVE_TOOL_ROOTS, so file tools stay contained until you configure roots. It does not write the key, live node addresses, or generated node JSON to the repository.
Double-click DaveLLM Launcher in ~/Applications for subsequent launches. Startup failures are written to ~/Library/Logs/DaveLLM/launcher.log. Stop any browser-mode process already using TCP port 8000 before launching the desktop app.
The launcher automatically discovers Homebrew whisper-cli. Local dictation uses the model under DAVE_DATA_DIR/models/ggml-tiny.en.bin; override either path with DAVE_WHISPER_BIN or DAVE_WHISPER_MODEL.
source venv/bin/activate
export DAVE_API_KEY='<local-secret>'
export DAVE_NODES='[...]'
uvicorn app:app --host 127.0.0.1 --port 8000Open http://127.0.0.1:8000. Browser mode prompts for the key and keeps it in sessionStorage; it is not written to persistent localStorage. /health is public. Protected endpoints fail with 503 when the server key is unset and 401 when the supplied key is missing or wrong.
The UI loads nodes from GET /nodes, then loads the selected node's inventory from GET /nodes/{node_id}/models. The router queries Ollama GET /api/tags. Chat requests are accepted only when node_id is a configured node and model appears in the loaded inventory for that node. Plain chat (POST /chat, POST /chat/stream, and conversation-summary generation) uses Ollama's native POST /api/chat through the shared helper in davellm_ollama.py, which sends the router's own num_ctx (the model's window capped at DAVE_CHAT_NUM_CTX, default 16384) and keep_alive (DAVE_CHAT_KEEP_ALIVE, default 30m, for chats in a conversation), so a node's own default context no longer decides plain chat. The bounded tool loop (POST /tools/agent/run and the lifecycle run routes) also uses native POST /api/chat (DL-TRANSPORT-01b): it sends the tool schemas, the same num_ctx and keep_alive, and a byte-capped read; DaveHarness reads the native reply directly, and only the tool-call arguments and tool-result names it writes back are converted. Ollama's OpenAI-compatible endpoint ignores num_ctx and keep_alive (verified on Ollama 0.33.3). Sampling is unchanged: /v1 sent top_p 1.0, so the native payload sends top_p 1.0 too instead of the model's Modelfile default. Request-field mapping is in INTEGRATION.md.
A DAVE_NODES entry can carry an optional capability profile (DL-ROUTE-01): "profile": {"compute": "cpu", "prompt_token_limit": 1300, "model_prompt_token_limits": {"gpt-oss:120b": 2500}}. prompt_token_limit is the prompt size, in estimated tokens for everything sent (system prompt, project context, history and the new message), that the node reads in reasonable time. A per-model entry overrides it. GET /nodes returns the profile, and the node picker labels a CPU-only node "(CPU, slower)". When a streamed prompt is over its node's limit, the reply bubble says so while it waits (DL-ROUTE-02). The bubble also says whether the model is loading into memory or already loaded (DL-ROUTE-03, from the node's /api/ps), and whether the node is busy with other replies (DL-ROUTE-04). The limits are advisory: the request still goes to the node and model you picked. The launcher sets Dominic (CPU) to 1,300 tokens and gpt-oss:120b on Duncan to 2,500, from the 2026-09-25 benchmarks. Every prompt carries the ~900-token default system prompt, so on Dominic most conversations show the warning. A malformed profile is ignored with a startup warning, and the node stays registered.
While a streamed reply is still silent, the chat shows what the model is doing: "Waiting for the model (loading it and reading your message)…", then "Thinking…" if it reasons first. When the reply finishes, a small line under it gives the speed, e.g. 58.2 tok/s · first token 1.9 s · 256 tokens · model load 1.5 s. Model load is shown only when it took at least 1 s. These are the status and stats stream events (DL-UX-01) in INTEGRATION.md. The speed line is not saved with the conversation.
The repository intentionally contains no real node addresses or verified model inventory. Cluster reachability and installed models remain runtime-dependent.
GET /nodes/{node_id}/models and GET /nodes/status both return an error field so an empty inventory explains itself instead of looking like an empty cluster:
error: nullwith an emptymodelslist means the node answered/api/tagsand has nothing pulled. Runollama pull <model>on that node.errorset means the node was not reached. The message names the cause: connection refused, an HTTP status, or a timeout.- A node that answers
curl http://127.0.0.1:11434/api/tagslocally but refuses the router is bound to loopback only. Start it withOLLAMA_HOST=0.0.0.0:11434. - A malformed
DAVE_NODESvalue registers zero nodes. The router now prints the parse error and the expected JSON shape at startup instead of failing silently.
DAVE_NODE_TIMEOUT sets the per-node /api/tags timeout in seconds and defaults to 10. Raise it for nodes that are slow to answer while loading a large model.
DAVE_DATA_DIR relocates all eight persistence files and the project-upload directory. When unset, the current working directory remains the default.
dave_conversations.jsondave_projects.jsondave_settings.jsondave_project_context.dbdave_vectors.dbfeedback.dbperformance.dbcost_log.jsonlproject_uploads/
Tools are disabled by default in the router; the macOS launcher turns them on. To enable them elsewhere, set DAVE_ENABLE_TOOLS=true and provide DAVE_TOOL_ROOTS as a JSON array of absolute paths. File read, write, and append operations share the same containment check. shell.exec remains disabled unless DAVE_ENABLE_SHELL_TOOL=true is also set. DAVE_ENABLE_EXTENDED_TOOLS=true, honored only when DAVE_ENABLE_TOOLS=true, adds the read-only file.list, file.search, file.read_lines, md.outline, md.section, git.status, git.diff, git.log, git.show, the run-scoped native tools project.notepad.read, project.brain.read, project.artifacts, chat.search, and cluster.status, and file.edit, which replaces exact text in an existing file only after you approve a before-and-after preview; the file tools hide protected paths such as .env files, keys, and SSH folders, and never follow symlinks out of a root. See docs/DAVELLM_TOOLS.md for their arguments and limits. web.fetch accepts only bounded public HTTP/HTTPS responses and validates DNS plus each redirect target.
With extended tools on, web.search searches the web through a self-hosted SearXNG instance named by DAVE_SEARCH_URL (its JSON format must be enabled), and web.read returns a public page as readable text. Together they let a Run tools request look something up and read the source. See docs/DAVELLM_TOOLS.md and docs/OPERATIONS.md.
The tool registry publishes one JSON schema per active tool. POST /tools/agent/run sends the current schemas to Ollama on every bounded model step, validates arguments, logs call and result timing, returns the complete transcript, stops after eight steps by default, and pauses before tools marked as requiring approval. The pending response includes a run_id plus exact-call metadata. POST /tools/agent/resume accepts that run ID, call ID, SHA-256 argument digest, and an approve or deny decision for up to 300 seconds. Approval executes the stored canonical arguments without replaying the paused model step; denial records an operator-denied tool result and continues.
The generic registry and bounded-loop implementation lives in the in-process daveharness package at version 1.0.0-rc.1. DaveLLM imports that public API and retains concrete tools, authentication, the native Ollama transport (davellm_ollama.py), persistence, and HTTP routes in app.py, with the native /api/chat transport in davellm_ollama.py. Root tool_executor.py and daveharness/executor.py remain compatibility re-exports; new code should import daveharness.
DaveHarness 0.3.0 adds CONTRACT_VERSION = 1 and the separate encode_contract/decode_contract envelope for ParsedToolCall, PendingCall, ToolExecution, and ExecutorOutcome. The envelope has exactly contract_version, contract_type, and payload; canonical JSON uses sorted keys, compact separators, and UTF-8 without ASCII escaping. Unknown versions, types, fields, invalid timestamps, non-finite numbers, and non-JSON nested values are rejected. DaveHarness 0.4.0 adds injected policy decisions, definition fingerprints, and immutable budgets. DaveHarness 0.5.0 adds versioned RunSnapshot, exact ApprovalDecision, and compare-and-swap resume_run/decide_run operations; see H3 state contract. DaveHarness 0.6.0 adds cancellation and deadline contracts. DaveHarness 0.7.0 adds metadata-only lifecycle events. DaveHarness 0.8.0 adds the bounded run store and instance-owned facade. DaveHarness 0.9.0 adds the DaveLLM lifecycle integration. The legacy route retains its original step and error limits. ExecutorOutcome.to_dict() and legacy /tools/agent/* responses keep their existing fields.
Tool handlers are synchronous by default. Coroutine handlers require both async_handler=True and a host-supplied name allowlist; DaveLLM permits web.fetch, web.search, and web.read. Definitions declare cancellation as bounded or abandon, and write or approval-required tools cannot use abandon. Existing tool status strings are unchanged. The additive termination field is completed, deadline_abandoned, denied, or error; deadline_abandoned means the response deadline elapsed and the underlying synchronous worker may still be running. Pending calls live only in the current process and do not survive restart.
The Chat header opens a layered instruction editor. It shows the global default, attached project instructions, session override, precedence, live character and token estimates, and the exact effective system text. Saves apply to the next message without restarting the application. Session overrides and the editable runtime global default each have a one-action reset.
Completed user and assistant messages expose keyboard-reachable Copy and Add to notepad actions. Copy preserves raw message source. The plain-text notepad persists per project, autosaves, stays inside Chat, can accept a selection from either message role, and can send its full contents as one user message.
Project Home is the inspectable owner for exactly four request-context components: Project Instructions, File Context Uploads, Artifact History, and BRAIN. Its baseline budget is deterministic: 25 percent instructions, 25 percent BRAIN, 30 percent files, and 20 percent artifacts. Unused capacity rolls forward to BRAIN, then files, then artifacts; instructions and protected BRAIN text are rejected instead of silently truncated. Preview context assembles the exact next-request messages without calling a model.
BRAIN stores pinned facts, active work, and compactable recent context in dave_project_context.db. Threshold and explicit compaction create immutable revisions; duplicate, resolved, superseded, and raw tool-log lines can be removed while pinned and active tiers remain verbatim. Delete is recoverable for DAVE_BRAIN_RECOVERY_DAYS, after which the daily worker permanently clears old content and starts a fresh revision history.
The authenticated local CLI uses the running router and DAVE_API_KEY:
python scripts/project_context_cli.py show <project-id>
python scripts/project_context_cli.py pin <project-id> --text 'Decision: verify before release.'
python scripts/project_context_cli.py compact <project-id>
python scripts/project_context_cli.py revisions <project-id>
python scripts/project_context_cli.py restore <project-id> 2Project-context configuration:
DAVE_MODEL_CONTEXT_DEFAULT— fallback model window, default32768.DAVE_MODEL_CONTEXT_WINDOWS— JSON object of model IDs to context-window tokens. For plain chat the router sends each window, capped atDAVE_CHAT_NUM_CTX, to Ollama asnum_ctx, and sizes plain chat's project-context budget (and the context preview) with the same capped number. Agent runs, which still use/v1and send nonum_ctx, budget with the uncapped window.DAVE_CHAT_NUM_CTX— largest context the router requests for plain chat, default16384; values below8192are raised to8192, the smallest window that leaves room for project context after the output and safety reserves.DAVE_CHAT_KEEP_ALIVE— how long Ollama keeps a model loaded after a chat in a conversation, default30m; empty uses the server default.DAVE_NODE_CONNECT_TIMEOUT— seconds to connect to a node before giving up, default10.DAVE_NODE_FIRST_CHUNK_TIMEOUT— seconds a streamed chat may wait for its first line (model load plus prompt reading), default300.DAVE_NODE_IDLE_TIMEOUT— seconds a streamed chat may wait between lines once tokens flow, default120.DAVE_NODE_TOTAL_TIMEOUT— seconds for a whole non-streaming reply (/chatand each tool-loop model turn), default600.DAVE_PROJECT_CONTEXT_TOKENS— new-project context budget, default16384.DAVE_BRAIN_COMPACT_TOKENS— new-project compaction threshold, default3072.DAVE_BRAIN_RECOVERY_DAYS— soft-delete recovery window, default30.
The Project Homepage uses a vendored GSAP 3.15.0 core timeline for its precision-control-deck reveal and context-preview feedback. It animates transform and opacity only, switches directly to final states under prefers-reduced-motion: reduce, and performs no CDN request. The vendored notice is in static/vendor/gsap/NOTICE.md.
The browser client derives at most three deterministic Suggested next actions from the current composer, attachment type, validated project/node/model selection, and response shape. Prediction generation lives in static/anticipation.js; it performs no network, DOM, or storage work, and suggestion chips never submit or change context without a visible user action. Valid last-used selections may be restored only on the empty startup state after inventory and node-health checks, with a visible status and immediate Undo. The client no longer calls /route/decision while sending a message, so the selected model remains under manual control.
Suggestion preferences use the versioned davellm_anticipation_v1 local-storage record. Its schema is limited to stable IDs, booleans, capped counters, and timestamps; prompt/response text, attachment names or contents, credentials, node URLs, and system prompts are excluded. Unsent composer text and the active Chat/History/Runtime mobile tab remain session-only. Reset Suggestions removes only the versioned suggestion record.
Install requirements-dev.txt in the project virtual environment before running these checks.
source venv/bin/activate
python -m py_compile app.py
python -m py_compile project_context.py scripts/project_context_cli.py
python -m py_compile tool_executor.py
python -m py_compile davellm_ollama.py
python -m py_compile davellm_web.py
python -m compileall -q daveharness
python -m mypy daveharness
python -m pytest -q
node --check static/app.js
node --check static/anticipation.js
node --check static/prompt-contract.js
node --check static/vendor/gsap/gsap.min.js
node --check desktop/main.js
node --check desktop/preload.js
bash -n deploy/check-cluster.sh scripts/verify-cluster.sh scripts/macos/install-launcher.sh scripts/macos/install-whisper-runtime.sh scripts/macos/launch-davellm.sh
npm ci
npm ls --depth=0
npm audit
git diff --checkRuntime UI files are under static/; only that directory is mounted at /. The separate docs/ content is published through GitHub Pages and is not used by the desktop runtime.
package.json pins Electron ^44.0.0, upgraded on 2026-08-26 to resolve the prior high-severity advisory set. Keep Electron on a supported major and rerun npm audit after dependency changes.
DaveHarness 1.0.0-rc.1 adds H8 offline qualification, bounded JSON admission, and Python 3.12–3.14 CI. See H8 qualification. This candidate does not establish live-model qualification or human acceptance. DaveLLM remains 2.1.0.
H9 action 55 adds the authorized live-evaluation runner scripts/evaluate_live_daveharness.py, which runs DaveLLM's Ollama adapter and file/system tools inside disposable roots and reports the action 56 thresholds per model. See H9 live evaluation. No target model is qualified yet, and DaveHarness remains 1.0.0-rc.1.