A measurement instrument and linter for prompts. Paste a system prompt, a CLAUDE.md, or a
set of agent instructions, and get back which lines actually carry behavioural weight and
which are ballast, measured empirically against a local model instead of argued from taste.
Every score comes with a credible interval, and it runs entirely on your machine.
It's ablation-led: to score a line, the tool removes it, re-runs a fixed suite of probe tasks, and measures how much the model's behaviour changed. Token attention is used only as a cheap pre-screen and an explainer, never as the verdict.
- Splits a prompt into discrete instructions.
- Scores each one 0 to 100 by ablation, and classifies it load-bearing, contributing, decorative, or contradicted.
- Puts a Bayesian credible interval on every score, and marks a line as noise when its effect can't be told apart from run-to-run variance, rather than dressing it up as a low number.
- Runs a phrasing duel: two wordings of the same instruction go head to head, and it names a winner only when the difference clears a region of practical equivalence.
- Two surfaces, same numbers: an interactive web app and a CLI.
- Python 3.11 or newer.
- A local model runtime. The easiest is Ollama: no gated downloads, no GPU wrangling, and it manages the models for you. A CUDA GPU helps but isn't required for the small models.
-
Install Ollama, then pull the three models it uses:
ollama pull qwen3.5:2b # subject: the model under test ollama pull qwen2.5:7b-instruct # judge: scores compliance (never judges itself) ollama pull nomic-embed-text # embedder: output-divergence signal
-
Install the package (editable, so the shipped probe suite is found in place):
pip install -e ".[web]" -
Run the web app:
proseweight web
Open http://127.0.0.1:8790, paste a prompt, and click Measure weights. The Duel tab pits two phrasings against each other. Both pages have Subject and Judge dropdowns: the subject list is populated from the models Ollama has pulled, plus the Anthropic frontier models (Opus, Sonnet, Haiku, Fable); the judge stays local.
Or measure straight from the command line:
proseweight scan your-prompt.txt --depth deep --probes 6 --n 3You get a terminal readout with a weight bar, a 95% credible interval, and a verdict per
line, plus a machine-readable report with --json report.json.
A deep audit runs a lot of model calls and takes a few minutes; the 7B judge is the pacing
item. Lower --probes and --n for a faster first look.
| Role | Default | Notes |
|---|---|---|
| Subject | qwen3.5:2b |
The model under test. A bigger subject (7B) discriminates better. |
| Judge | qwen2.5:7b-instruct |
A different model, so nothing judges its own output. |
| Embedder | nomic-embed-text |
Semantic distance between outputs. |
All swappable:
proseweight scan p.txt --subject qwen2.5:7b-instruct --judge qwen3.5:4bproseweight scan <file>measure a prompt and print the readout (--jsonfor the report)proseweight webthe interactive app (Scan and Duel)proseweight export <report.json> --html out.html [--png card.png]a self-contained reportproseweight diff <v1.json> <v2.json>weight changes between two versions of a promptproseweight lint <report.json> --baseline weights.jsonCI gate; exits non-zero on a load-bearing regression or a dead-weight budget breachproseweight servea small local HTTP API
For each instruction the tool removes it, re-runs the probe suite N times, and measures the behavioural delta against the full prompt. The delta blends three signals with fixed, documented weights: an LLM-judge rubric score, an embedding distance between outputs, and task-specific programmatic checks. The statistics are Bayesian throughout, computed in numpy: every weight is a posterior with a credible interval, and cross-instruction shrinkage does the job a multiple-comparisons correction would. Attention is a pre-screen that ranks which lines to ablate first, and an explainer under the "how it works" view. It never sets a score.
- Weights are specific to the probe suite and the model. There is no universal prompt score; the model-comparison view exists instead.
- Small models shift their output when you remove almost anything, so genuine dead weight is hard to surface without a fuller probe suite and a larger subject model.
- An optional frontier API judge (Anthropic, key from
ANTHROPIC_API_KEYonly) is available; runs that use it are flagged best-effort and waive same-seed reproducibility.
The weight linter asks which of your instructions do anything. CacheScope, a sibling that shares the same bench, asks which of your bytes cost you money for nothing. It captures your real Claude traffic locally, byte-diffs why each prompt-cache miss happened, maps it onto the real cache breakpoints, and prices the avoidable waste: your CRLF line endings cost you £4.12 this month, and here is the byte that did it.
proseweight cache lint <file>...static cache-hygiene lint with no captured data (CRLF drift, trailing whitespace, volatile headers, concatenation order), each finding an estimated monthly cost. Add--baselinefor the CI gate: it fails a commit that makes a file cache-hostile, with the cost in the message.proseweight cache servea local recording proxy. Point an app at it withANTHROPIC_BASE_URLand it captures the exact bytes and the reported usage.--diagnosticsopts into Anthropic's cache-diagnostics beta.proseweight cache ingest <transcript>reads Claude Code session transcripts for the subscription path;--reconstructrebuilds the growing prefix with a calibrated confidence.proseweight cache analyse/ledger/exportthe divergence analysis, the waste ledger with a headline figure, and a self-contained HTML report with a byte-level diff on every miss.
Everything stays on your machine, and the meter forks by how you pay: measured pounds on the pay-as-you-go API, quota plus a labelled shadow-price on a subscription (where a pound bill would be a lie). The cost engine is deliberately self-contained behind a versioned contract, so it sits cleanly between two sibling tools: OmnisRouter captures the traffic, CacheScope analyses it, and OmnisVigil reports on it.
You can profile a closed API model as the subject instead of a local one, either from the Subject dropdown in the web app or from the command line:
proseweight scan p.txt --backend anthropic --subject claude-opus-4-8 --judge qwen2.5:7b-instructAny claude-* id works (Opus, Sonnet, Haiku, Fable). Only the subject's generations hit the
paid API; the judge and embedder stay local, and the judge being a different model avoids
self-judging. Three caveats: frontier models don't expose attention, so there's no attention
view for these runs; they aren't seed-deterministic, so the run is flagged best-effort and
same-seed reproducibility is waived; and they cost money, so a deep audit of a large prompt
adds up. The key is read from ANTHROPIC_API_KEY in the environment only, and the CLI prints
the token usage at the end so you can see what a run cost. The smaller models (Haiku, Fable)
are far cheaper for a first pass.
pip install -e ".[runtime]"
proseweight scan p.txt --backend hf --subject Qwen/Qwen2.5-1.5B-InstructThis downloads the weights from the Hugging Face Hub and, realistically, wants a GPU.
The original single-file attention demo (attnscope / attnduel) still lives in
prose_weight_visualiser.py. Its attention view became the "how it works" layer here, and
its two-prompt fight became the phrasing duel.
pip install -e ".[dev]"
pytest # the deterministic engine is covered without any model runtime
ruff check .