A friendly guide for developers and curious users who want to evaluate large language models (LLMs) running locally via Ollama.
Last updated: May 2026 · Version: 0.1.0
- What this tool does
- Concepts in plain English
- Installation prerequisites
- Your first evaluation in 5 minutes
- Using the web UI
- Using the command line
- Writing your own test suites
- Using public benchmarks (MMLU, GSM8K, etc.)
- Running remotely via SSH
- Reading the reports
- Troubleshooting
- Command reference
- Glossary
You have one or more LLMs running locally through Ollama (for example
qwen3.6:27b or llama3:8b). You want to know:
- How fast is each model?
- How accurate is it on questions you care about?
- Which model should I pick for my task?
- Has a newer model improved on the old one?
This tool runs a set of questions through each model, scores the answers, measures the speed, and gives you a report — both a spreadsheet-friendly JSON file and a human-readable Markdown summary.
You can drive it three ways:
- A web UI in your browser (easiest to start with).
- A command-line tool (best for scripting and CI pipelines).
- An HTTP API (if you're building your own dashboard on top).
Everything runs on your own hardware. Nothing is sent to a cloud service.
"Is
llama3:8bfaster thanqwen3.6:27bon basic arithmetic, and which one gets more answers right?"
The tool runs the same arithmetic questions on both models, captures the response, checks if the response contains the right answer, measures how many tokens per second each model produced, and gives you a side-by-side comparison.
Before you touch anything, here are the words used throughout this manual.
A program that runs LLMs on your computer. It listens on port 11434 and
exposes a simple HTTP API. You pull models (like downloading apps) and then
ask them questions. Install from ollama.com/download.
A specific LLM identified by a name and tag, like llama3:8b or
qwen3.6:27b. The part after the colon usually indicates size or variant.
One question (called a "prompt") plus the expected answer (or a rule for what counts as a correct answer) plus a name. Example:
id: arithmetic-simple
prompt: "What is 2 + 2?"
expected_output: "4"
metric: contains "4"
A named collection of test cases. You typically group related questions together — a "math" suite, a "code" suite, a "safety" suite.
A rule that scores a response. Built-in ones:
- exact-match — response must equal the expected output letter-for-letter.
- regex-match — response must match a pattern (like "any 4 in the text").
- contains — response must contain certain substrings.
- json-schema-valid — response must be valid JSON matching a shape.
- length-range — response must be within a certain length.
- llm-as-judge — another LLM scores the response against a rubric.
- response-capture — always passes; records the raw response for later grading.
One execution of the tool. A Run takes the configuration ("evaluate these models against these suites") and produces a Run Report.
The output of a Run. Contains every response, every metric score, timings,
and totals. Written as both report.json and report.md.
How many times each test case is executed per model. Useful for measuring variance in non-deterministic models. Default is 1.
How many questions can be in flight at the same time. Default is 1 (sequential) to keep GPU memory usage predictable.
The backend is the Python program that orchestrates everything. The UI is a
React web app that talks to the backend. When you run serve, the backend
starts a local web server and also serves the UI bundle so you can open it
in a browser.
You need four things:
| What | Version | Why |
|---|---|---|
| Python | 3.11 or newer | Backend language |
| Node.js | 18 or newer | Builds the web UI bundle |
| Ollama | 0.5 or newer | Serves the LLMs being evaluated |
| A small Ollama model | any | Something to evaluate |
On Linux:
curl -fsSL https://ollama.com/install.sh | shOn macOS or Windows, download from ollama.com/download.
After install, pull at least one model:
ollama pull llama3:8b # 4.7 GB, general-purpose
# or, if you have more RAM / GPU memory
ollama pull qwen3.6:27b # ~17 GB, reasoning modelVerify Ollama is running:
curl http://localhost:11434/api/version
# → {"version":"0.20.3"}You have two options:
The repo ships with a one-command install script that handles Python venv, backend install, UI build, schema regeneration, and a smoke test.
Linux/macOS:
./scripts/install.shWindows PowerShell:
.\scripts\install.ps1Both installers are idempotent — run them again after git pull to pick up
dependency changes. Pass --skip-tests (bash) or -SkipTests (PowerShell) to
skip the unit-test run and speed up re-installs.
This tool is not yet on PyPI or npm. You clone the repository and install from source.
# 1. Get the code (example path; put it wherever you like)
git clone <your-repo-url> AI-Model-Evaluation
cd AI-Model-Evaluation
# 2. Python environment
python3 -m venv .venv
source .venv/bin/activate # On Windows PowerShell: .venv\Scripts\Activate.ps1
pip install -U pip
pip install -e "backend[dev]" # Installs the backend editable
# 3. UI dependencies and bundle
cd ui
npm install # Downloads ~200 packages
npm run build # Produces ui/dist/
cd ..Verify either install:
python -m ollama_evaluator.cli --help
# You should see a help screen listing subcommandsThe repository ships with an examples/ directory containing a ready-to-run
suite. There are two ways to get going: the one-button launcher that starts
every service for you, or the step-by-step commands below.
./scripts/start.sh # Linux / macOS.\scripts\start.ps1 # Windows PowerShellOr just double-click start.bat on Windows.
What it does, in order:
- If
.venv/orui/dist/are missing, runsinstall.sh/install.ps1for you. - If Ollama isn't already running on port 11434 and you have the
ollamaCLI installed, startsollama servein the background. - Starts the FastAPI backend on port 8765, with the built UI mounted at
/. - Waits until every service responds to a health check before handing back.
- Tails the backend log; Ctrl-C stops every service it started (but not an Ollama it found already running — that one keeps going so other apps on your machine aren't disturbed).
After a few seconds you'll see:
All services ready.
Web UI http://localhost:8765/
Health probe http://localhost:8765/api/health
REST API docs http://localhost:8765/openapi.json
Open the UI URL in your browser.
Stop everything later with:
./scripts/stop.sh # or Ctrl-C if start.sh is still in the foreground.\scripts\stop.ps1Useful flags:
--background/-Detach— return immediately instead of tailing the log.--port 9000/-BackendPort 9000— pick a different backend port.--dev/-Dev— also start the Vite dev server with hot-reload.--skip-ollama/-SkipOllama— skip the Ollama check/start.--host 127.0.0.1— bind only to localhost (default is 0.0.0.0 = LAN).
Open examples/suites/reasoning-basics.yaml. It has three questions:
name: reasoning-basics
defaults:
temperature: 0.0
max_tokens: 512
test_cases:
- id: arithmetic-simple
prompt: "What is 2 + 2? Answer with just the number."
expected_output: "4"
metrics:
- name: regex-match
params:
pattern: "\\b4\\b" # "look for a standalone 4 anywhere"
- id: capital-of-france
prompt: "What is the capital of France? Answer with just the city name."
expected_output: "Paris"
metrics:
- name: contains
params:
substrings: ["Paris"]
mode: any
- id: python-syntax
prompt: "Write a Python one-liner that returns the square of 5. Respond with code only."
metrics:
- name: regex-match
params:
pattern: "25"Open examples/config.qwen.yaml:
ollama_base_url: http://localhost:11434
suites_dir: examples/suites
output_dir: examples/runs
run:
models:
- qwen3.6:27b # ← change this to your model
suites:
- reasoning-basics
repetitions: 1
concurrency: 1
ollama_timeout_s: 300.0Change models: to whatever you pulled, e.g. - llama3:8b.
python -m ollama_evaluator.cli validate-suite examples/suites/reasoning-basics.yaml
# → OK: reasoning-basics (3 test cases)If the YAML has typos, this step tells you the exact field and line.
python -m ollama_evaluator.cli list-models
# → llama3:8b sha256:abc... 8BThis confirms the backend can reach Ollama.
python -m ollama_evaluator.cli --config examples/config.qwen.yaml runYou will see HTTP log lines, then:
run f3c9a1...: 3 executions, passed=2, failed=1, error=0, timeout=0
The process exits with code 0 if every test passed, 1 if any failed, 2 on setup problems like "Ollama not reachable". This matters for CI pipelines.
cat examples/runs/f3c9a1.../report.mdYou'll see per-model and per-suite tables. For a deeper look:
python3 scripts/remote_show_report.pyThis prints every response plus timings.
If you prefer a browser experience, start the server:
OLLAMA_EVAL_UI_DIR=$(pwd)/ui/dist \
python -m ollama_evaluator.cli --config examples/config.qwen.yaml serve \
--host 0.0.0.0 --port 8765Flag explanations:
--host 0.0.0.0— listen on every network interface (use127.0.0.1to bind only to the local machine).--port 8765— TCP port to listen on. Change if something else uses this.OLLAMA_EVAL_UI_DIR— tells the backend where the built UI bundle lives so/serves the React app.
Leave this command running. Open a browser to http://<host>:8765/.
-
New Run (
/runs/new). Pick models (chips with size / quantisation shown in the hover tooltip) and pick suites (grouped by category — Reasoning, Knowledge, Coding, Math, Instruction, Multilingual, Long context, Safety, Open-ended). Each suite shows name · case count · rough ETA inline; hover reveals the full methodology caveats. Set repetitions, concurrency, optional tag filter, then click Submit Run. You are immediately navigated to the Run Detail page. -
Run Detail (
/runs/:id). Live progress bar, four KPI cards (passed / failed / errored / timeout), a live per-model table updating as each test case completes, and an execution table with status pills. A red banner appears if the stream disconnects. There is a Cancel button while the Run is in progress. When the Run finishes, the report panel shows Summary by model, a Model × Suite breakdown table, every per-test-case row, and download links for the JSON and Markdown reports. -
History (
/history). Every past Run. Filter by model, suite, status, or date (URL-encoded so links are shareable). Check two rows and click Compare selected to open a diff view. -
Compare (
/compare?a=…&b=…). Side-by-side metric scores and performance numbers for two Runs, with signed deltas coloured green (B better) or red (A better).
The top-right Light / Dark toggle flips the theme instantly. First
visit auto-matches your OS's prefers-color-scheme; subsequent visits
respect your stored choice. No flash of wrong theme on reload.
Stop the server with Ctrl-C.
For everyday use, the CLI is faster than clicking through the UI.
# Activate the venv every new terminal session
cd AI-Model-Evaluation
source .venv/bin/activate
# See help
python -m ollama_evaluator.cli --help
python -m ollama_evaluator.cli run --help
# Validate a suite file offline (does not need Ollama)
python -m ollama_evaluator.cli validate-suite path/to/suite.yaml
# List what Ollama has available
python -m ollama_evaluator.cli list-models
# Execute a Run
python -m ollama_evaluator.cli --config examples/config.qwen.yaml run
# Compare two Runs by id (ids come from the output of `run`)
python -m ollama_evaluator.cli --config examples/config.qwen.yaml \
compare f3c9a1... d712fe...
# Start the HTTP + WebSocket server
python -m ollama_evaluator.cli --config examples/config.qwen.yaml serve| Flag | Default | Purpose |
|---|---|---|
--config <path> |
required | Config file path |
--output-dir <path> |
from config | Override output directory |
--log-level {debug,info,warn,error} |
info |
Verbosity |
--dataset-mode {local,remote} |
local |
Local cache vs HuggingFace Hub |
--hf-cache-dir <path> |
HF default | Where HuggingFace caches datasets |
| Code | Meaning |
|---|---|
0 |
Every test case passed |
1 |
At least one test case failed |
2 |
Preflight error (Ollama unreachable, missing model, etc.) |
This matters when scripting. In a CI pipeline you can do:
python -m ollama_evaluator.cli --config ci.yaml run
if [ $? -ne 0 ]; then
echo "Evaluation regressed, failing build"
exit 1
fiA suite is a YAML (or JSON) file in your suites_dir. Minimum shape:
name: my-first-suite
test_cases:
- id: hello-world
prompt: "Say hello"
metrics:
- name: contains
params:
substrings: ["hello", "Hello"]
mode: anyThe id must be unique within the suite. prompt and metrics are required.
Everything else is optional.
name: math-deep
version: "1.0"
description: "Math problems of increasing difficulty"
defaults:
temperature: 0.0 # Lower = more deterministic
max_tokens: 256
stop_sequences: []
test_cases:
- id: add-tiny
prompt: "What is 7 plus 5?"
expected_output: "12"
tags: [math, easy]
metrics:
- name: regex-match
params:
pattern: "\\b12\\b"
- id: mul-tiny
prompt: "What is 7 times 5? Answer with just the number."
expected_output: "35"
tags: [math, easy]
metrics:
- name: regex-match
params:
pattern: "\\b35\\b"
- id: word-problem
prompt: |
A train leaves station A at 2pm going 60mph.
Another train leaves station B at 3pm going 80mph.
They are 340 miles apart. When do they meet?
Answer in the form "HH:MM pm".
expected_output: "5:00 pm"
tags: [math, hard, word-problem]
temperature: 0.2 # Per-case override
metrics:
- name: contains
params:
substrings: ["5:00 pm", "5 pm", "17:00"]
mode: anysystem_prompt— string prepended to the request as a system message.expected_output— reference answer, used byexact-match,regex-match.reference_data— arbitrary structured data (used by external graders).tags— list of strings, for filtering at run-time (tag_filter:).temperature,max_tokens,stop_sequences— per-case generation overrides.
In your config:
run:
suites: [math-deep]
tag_filter: [easy] # Only test cases whose tags include "easy"An empty tag_filter: [] means "run every test case".
- Keep prompts short and specific. Ambiguous prompts make metrics unreliable.
- Use
temperature: 0.0unless you want randomness. For reproducible scoring, determinism is your friend. - Write forgiving metrics.
containsorregex-matchusually works better thanexact-matchbecause models often add whitespace or explanations you don't care about. - Give reasoning models enough
max_tokens. Qwen and similar models emit a chain-of-thought; if you cap them at 64 tokens they'll never get to the actual answer. 512–1024 is a safe starting point. - Version your suites with Git. Commit the YAML so you can reproduce any historical result.
The tool ships with adapters for five well-known benchmarks and a generic HuggingFace loader.
We ship 33 suites totalling ~4 910 cases. The hand-authored ones
(reasoning-basics, factual-qa, instruction-following, json-output,
math-word-problems, code-generation-basics, reasoning-advanced,
safety-refusal, multilingual-basic, long-context-probe,
llm-as-judge-general, python-bugfix-mini, shell-bash-basics) live
under examples/suites/. The benchmark-backed ones were materialised from
HuggingFace via helper scripts under scripts/:
| Benchmark | Dataset | Metric | Notes |
|---|---|---|---|
| MMLU | cais/mmlu |
regex-match (letter A/B/C/D) | 57-subject academic MCQ |
| HellaSwag | Rowan/hellaswag |
regex-match | Commonsense sentence completion |
| TruthfulQA | truthful_qa:multiple_choice |
regex-match | MC1 form |
| GSM8K | openai/gsm8k |
regex-match + numeric equality | Grade-school math |
| HumanEval | openai_humaneval |
response-capture | Capture-only; execution-based pass@1 grading pending |
| BigBench-Hard | lukaemon/bbh (8 subsets) |
regex-match on Final answer: |
Hard reasoning |
| ARC Challenge / Easy | allenai/ai2_arc |
regex-match | Science MCQ, two difficulty splits |
| PIQA | piqa:plain_text |
regex-match (A/B) | Physical commonsense |
| WinoGrande | winogrande:winogrande_xl |
regex-match (A/B) | Pronoun resolution |
| C-Eval | ceval/ceval-exam (6 subjects) |
regex-match (A/B/C/D) | Chinese academic MCQ |
| MATH-500 | HuggingFaceH4/MATH-500 |
response-capture | Competition math |
| MBPP | google-research-datasets/mbpp |
response-capture | Python function writing |
| SQuAD v2 | rajpurkar/squad_v2 |
contains | Extractive reading + unanswerable |
| IFEval | google/IFEval |
response-capture | Instruction following |
| MT-Bench | HuggingFaceH4/mt_bench_prompts |
llm-as-judge | Open-ended, judge-scored |
| PubMedQA | qiaojin/PubMedQA:pqa_labeled |
regex-match (yes/no/maybe) | Biomedical QA |
| CRUXEval-in / out | cruxeval-org/cruxeval |
regex-match | Code-reasoning without sandbox |
| Spider | spider |
regex-match | Natural-language → SQL |
Let's say you want MMLU's "abstract algebra" subject:
python -m ollama_evaluator.cli convert mmlu \
--source ./cache/mmlu \
--output ./examples/suites \
--subjects abstract_algebraThis reads the dataset from the local cache and writes a normal YAML suite
under ./examples/suites/mmlu-abstract_algebra.yaml. From then on it behaves
like any hand-written suite.
You can target any HF dataset with a field map. Create squad-fields.yaml:
prompt: "question"
expected_output: "answers.text[0]"
tags_from: ["category"]Then:
python -m ollama_evaluator.cli convert hf \
--hf-ref "squad:plain_text:validation" \
--field-map squad-fields.yaml \
--name squad-small \
--limit 100 \
--output ./examples/suitesThe --limit sub-samples deterministically (with optional --seed).
Two modes:
local(default, recommended). Reads pre-cached JSONL/Parquet from disk. Zero network at run-time. Run theconvertcommand once; every future evaluation uses the cached file.remote. Streams rows from HuggingFace Hub at run-time. Requires internet. First access pulls and caches the dataset. Slower first run, always up-to-date.
Set globally in the config (dataset_mode: local) or per-run with
--dataset-mode.
If your workstation can't run large models (no GPU, not enough RAM), run the tool on a server with the hardware and drive it over SSH.
The repo ships with a script that tars the project, pushes it to the remote,
and runs install.sh there — all in a single command.
From Linux/macOS:
./scripts/deploy-remote.sh user@host [/target/dir]From Windows:
.\scripts\deploy-remote.ps1 -Target user@host -RemoteDir /home/you/evalUseful flags:
--skip-tests/-SkipTests— skip the remote unit-test run.--skip-ui/-SkipUI— backend only.--serve PORT/-ServePort N— start the server in the background after install.--key PATH/-Key PATH— use a specific SSH key.--no-install/-NoInstall— only sync files, don't runinstall.sh.
Example one-liner to deploy, install, and launch the web UI on port 8765:
./scripts/deploy-remote.sh --serve 8765 \
--key ~/.ssh/id_ed25519 \
user@host /home/user/ollama-evaluatorThen browse to http://<host>:8765/.
If you prefer to understand each step, here is what the script does under the hood.
Replace user@server with your values.
# On your local machine
ssh-keygen -t ed25519 -f ~/.ssh/id_ed25519_eval -N ""
ssh-copy-id -i ~/.ssh/id_ed25519_eval user@server # enter password onceVerify:
ssh -i ~/.ssh/id_ed25519_eval user@server "echo OK"
# → OKWindows PowerShell equivalent:
# Generate
ssh-keygen -t ed25519 -f "$HOME\.ssh\id_ed25519_eval" -N '""'
# Install the public key (PuTTY's plink works well here)
# After the password prompt, future commands need no password.
$pub = Get-Content "$HOME\.ssh\id_ed25519_eval.pub" -Raw
ssh user@server "umask 077; mkdir -p ~/.ssh; echo '$($pub.Trim())' >> ~/.ssh/authorized_keys; chmod 600 ~/.ssh/authorized_keys"# Build a clean tarball (excludes caches, node_modules, dist)
tar --exclude='__pycache__' --exclude='*.egg-info' --exclude='.venv' \
--exclude='node_modules' --exclude='dist' --exclude='.pytest_cache' \
--exclude='.mypy_cache' --exclude='.ruff_cache' --exclude='.hypothesis' \
-czf /tmp/eval.tgz backend ui shared examples README.md
scp -i ~/.ssh/id_ed25519_eval /tmp/eval.tgz user@server:/tmp/
ssh -i ~/.ssh/id_ed25519_eval user@server \
"mkdir -p ~/workspaces/AI-Model-Evaluation && \
cd ~/workspaces/AI-Model-Evaluation && \
tar -xzf /tmp/eval.tgz"ssh -i ~/.ssh/id_ed25519_eval user@server bash -s <<'EOF'
set -e
cd ~/workspaces/AI-Model-Evaluation
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -e "backend[dev]"
cd ui
npm install
npm run build
EOF# Run tests
ssh -i ~/.ssh/id_ed25519_eval user@server \
"cd ~/workspaces/AI-Model-Evaluation && source .venv/bin/activate && \
cd backend && python -m pytest -q"
# Trigger an evaluation
ssh -i ~/.ssh/id_ed25519_eval user@server \
"cd ~/workspaces/AI-Model-Evaluation && source .venv/bin/activate && \
python -m ollama_evaluator.cli --config examples/config.qwen.yaml run"Start the server on the remote and expose it to your LAN:
ssh -i ~/.ssh/id_ed25519_eval user@server \
"cd ~/workspaces/AI-Model-Evaluation && source .venv/bin/activate && \
OLLAMA_EVAL_UI_DIR=\$PWD/ui/dist \
nohup python -m ollama_evaluator.cli --config examples/config.qwen.yaml \
serve --host 0.0.0.0 --port 8765 > serve.log 2>&1 &"The nohup ... & keeps it running after you log out. Stop it with:
ssh user@server "pkill -f 'ollama_evaluator.cli.*serve'"Then open http://<server-ip>:8765/ in your browser.
Rather than re-uploading the full tarball each time, push only what changed:
# Single file
scp -i ~/.ssh/id_ed25519_eval backend/src/ollama_evaluator/api/app.py \
user@server:~/workspaces/AI-Model-Evaluation/backend/src/ollama_evaluator/api/app.py
# UI changed? Rebuild on the remote
ssh -i ~/.ssh/id_ed25519_eval user@server \
"cd ~/workspaces/AI-Model-Evaluation/ui && npm run build"Every Run produces three artifacts under <output_dir>/<run_id>/:
report.json— machine-readable. Every field from every test case.report.md— human-readable Markdown tables.../history.db— SQLite database indexing every past Run.
Open examples/runs/<run_id>/report.md:
# Run report: f3c9a1...
- **Status**: completed
- **Backend version**: 0.1.0
- **Ollama version**: 0.20.3
- **Started at**: 2026-05-10T06:04:35+00:00
- **Ended at**: 2026-05-10T06:05:49+00:00
## Models
- **qwen3.6:27b** (digest=sha256:a50..., parameter_size=27.8B)
## Suites
- **reasoning-basics**
## Per-model results
| Model | Passed | Failed | Mean tokens/s | Mean total ms |
|---|---|---|---|---|
| qwen3.6:27b | 2 | 1 | 11.52 | 24666.00 |
## Per-suite results
...
## Error summary
None.Each test case is a record like:
{
"model": "qwen3.6:27b",
"suite": "reasoning-basics",
"test_case_id": "arithmetic-simple",
"repetition": 1,
"status": "pass",
"response": "4",
"error_message": null,
"performance": {
"ttft_ms": 278.9,
"total_ms": 14066.0,
"prompt_tokens": 15,
"response_tokens": 162,
"tokens_per_second": 11.52
},
"metrics": [
{"name": "regex-match", "score": 1.0, "passed": true, "threshold": 1.0}
]
}- status —
pass,fail,error, ortimeout.pass/fail: the metrics ran and scored; result is just the score.error: something broke (network, HTTP 5xx, retry exhausted).timeout: Ollama took longer thanollama_timeout_s.
- ttft_ms — milliseconds to first streamed token. A latency proxy.
- total_ms — wall-clock time for the full response.
- response_tokens — tokens the model emitted.
- tokens_per_second — derived:
response_tokens / (total_ms / 1000). Higher is faster.
A test case has status = pass if and only if every metric reported
passed = true. A single failing metric turns the whole case into fail.
Metrics that raised during scoring (e.g. malformed regex) are recorded with
error set, but do not fail the case — other metrics still apply.
python -m ollama_evaluator.cli --config examples/config.qwen.yaml \
compare <old-run-id> <new-run-id>Output is a JSON ComparisonReport with:
- metric_diffs — per
(model, metric)present in both: mean score A, mean score B, signed difference. - performance_diffs — per model: mean tokens/sec A/B, mean total_ms A/B, and the differences.
In the UI, go to /history, check two rows, click Compare.
The backend can't reach Ollama at the configured ollama_base_url.
- Is Ollama running?
curl http://localhost:11434/api/version - Is your config pointing at the right host? For remote Ollama, set
ollama_base_url: http://<host>:11434. - Firewall blocking 11434?
The requested model isn't on the Ollama server.
# Check what's available
ollama list
# Pull the missing one
ollama pull qwen3.6:27b
# Or set `pull_missing_models: true` in your configYou're probably using a reasoning model with a small max_tokens. Reasoning
models like Qwen emit a "thinking" trace before the final answer. Increase
max_tokens to 512 or more in the suite's defaults:.
Run:
python backend/scripts/regen_schemas.pyThis usually means the Pydantic version changed and the committed OpenAPI / JSON Schema files need refreshing.
You're running an old version. This was fixed by moving the store open into the FastAPI lifespan. Pull the latest code and re-install.
- Is the browser hitting the right backend? Check the URL bar.
- Does
curl http://<host>:8765/api/modelsreturn a list? If yes, it's a UI bug; open the browser's dev console to look for errors.
# Who's on 8765?
ss -tlnp | grep 8765
# Kill the old server
pkill -f 'ollama_evaluator.cli.*serve'
# Or use a different port
python -m ollama_evaluator.cli --config X run --port 8766- First run of a model is always slow because Ollama loads weights into VRAM. Subsequent runs are much faster.
- Reasoning models spend most of their time thinking, not generating. Check
tokens_per_secondin the report — if it's close to the raw model speed, you're not bottlenecked. - Increase
concurrencycautiously. Higher concurrency reduces wall-clock time but needs more VRAM and can cause out-of-memory errors. A safe starting point is 1–2 for 27B+ models, 4–8 for 8B.
The tool tells you the exact file, test-case id, and field. Fix the YAML and
run validate-suite to confirm.
The examples/runs/ directory grows one subdirectory per run. To clean up:
# List runs from the CLI
python -m ollama_evaluator.cli --config examples/config.qwen.yaml \
compare --help # CLI doesn't have a prune command yet
# For now, delete manually:
rm -rf examples/runs/<run-id>Or use the REST API:
curl -X DELETE http://localhost:8765/api/runs/<run-id>Usage:
python -m ollama_evaluator.cli [GLOBAL OPTIONS] COMMAND [ARGS]
Global options:
--config PATH YAML or JSON config file
--output-dir PATH Override config.output_dir
--log-level debug|info|warn|error
--dataset-mode local|remote
--hf-cache-dir PATH
Commands:
list-models List models available on the Ollama server
run Execute a Run using --config
compare RUN_A RUN_B Compare two historical Runs
validate-suite FILE Validate a suite file offline (no Ollama)
serve [--host H] [--port P] Start the HTTP + WebSocket + UI server
convert mmlu|hellaswag|... Materialise a public benchmark to YAML
convert hf --hf-ref REF --field-map FILE --output DIR --name N [--limit N]
HTTP API endpoints (when `serve` is running):
GET /api/health Liveness
GET /api/models List models
GET /api/suites List suite names
GET /api/suites/{name} Full suite content
POST /api/runs Submit a Run (body = RunConfig)
GET /api/runs List persisted runs
GET /api/runs/{id} Run report
GET /api/runs/{id}/report.md Markdown report
DELETE /api/runs/{id} Delete a run
POST /api/runs/{id}/cancel Cancel an in-progress Run
GET /api/compare?a=ID&b=ID Comparison report
GET /openapi.json OpenAPI 3.1 document
WebSocket:
WS /api/runs/{id}/events Live event stream
| Term | Meaning |
|---|---|
| Ollama | Local LLM runtime. Listens on http://localhost:11434. |
| LLM | Large Language Model. |
| TTFT | Time-To-First-Token. The latency before any output arrives. |
| Tokens/s | Throughput in tokens per second. Higher = faster generation. |
| System prompt | Instructions prepended to the request that tell the model how to behave. |
| Repetition | A single execution of (model, test case). Multiple repetitions measure variance. |
| Concurrency | How many generate calls run in parallel. |
| Preflight | Checks the tool does before starting a Run: reachable Ollama, model present, suites valid. |
| Run id | Unique identifier, like f3c9a1b42e8b4.... Used as the primary key in the history database. |
| History store | Local SQLite database at <output_dir>/history.db. Indexes every Run you've executed. |
| Backend | The Python program that orchestrates Runs and serves the API. |
| UI | The React web app. |
| REST / WebSocket | Two protocols the backend speaks. REST for request/response; WebSocket for the live event stream. |
| OpenAPI | Machine-readable API description at /openapi.json. Used to generate the UI's typed client. |
- For unexpected failures, run with
--log-level debugand share the output. - For missing features, open an issue or PR.
- Requirements, design, and implementation tasks live under
.kiro/specs/ollama-model-evaluator/— a good place to see why the tool works the way it does.