Skip to content

Run anya against real models, and fix the twelve things that broke - #1

Draft
Karthik777 wants to merge 15 commits into
claude/anya-ml-tools-65pa0efrom
cursor/e2e-real-model-tests-bc10
Draft

Karthik777 wants to merge 15 commits into
claude/anya-ml-tools-65pa0efrom
cursor/e2e-real-model-tests-bc10

Conversation

@Karthik777

@Karthik777 Karthik777 commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

The tiny graphs in nbs/fixtures prove the plumbing. They say nothing about what a real export declares, and it turns out most real exports declare something anya read wrong. evals/e2e.py downloads eight repos, runs them over five COCO and Ultralytics photographs plus three generated sounds, and asserts on the answers. Everything below was found by that script failing.

The one that mattered most

Three of the four ONNX repos tested declare their input as ['batch_size', 'num_channels', 'height', 'width'] — every axis symbolic, and named. anya read the channel axis from the sizes, found no integers, guessed NHWC, and fed channels-last pixels into a channels-first graph:

Non-zero status code returned while running Conv node. C: 256 kernel channels: 3

chan_axis now reads the names when the sizes cannot say. This is the commonest shape optimum emits, so it is the difference between anya working on the Hub and not.

The rest

What Was Is
webnn/yolov8n size 224 from preprocessor_config.json, into a graph fixed at 640 the size the graph states wins; a config cannot override it
Xenova/segformer-… task classify, because the output is named logits segment; logits is what both call their output and so decides nothing
SegFormer's 150 classes on a 128×128 grid read as 128 channels of a 128×150 image channels-first, and the mask goes back onto the original picture
Xenova/dinov2-small a 98,688-long "embedding" 257 tokens of 384 mean-pooled to one vector
timm crop_size 224 with size 256 a 256 crop a 224 crop out of a 256 short side, timm's crop_pct
resample always bilinear as the config asks; bicubic moved top-1 on one of five photos
yolo box units assumed pixels read off the range: two of three yolo exports ship 0..1
SSD boxes normalised against the padded input, then scaled as if it were not through scale_boxes, like every other detector
ssdlite320's twelve raw heads decoded to [] raises at load, naming task=
yamnet's 521 sigmoid scores softmaxed a second time, so silence read 0.004 0.801, because being in 0..1 is the test, not summing to 1
yamnet's per-frame output flattened, so the argmax crossed frames frames averaged, which is what yamnet documents
labels.txt at the repo root, weights under onnx/ class_15 for a bus its 80 COCO names
--conf=0.6 '0.6' into a numpy comparison 0.6
a tool that raises a traceback where a chat expects JSON {"error": …}, and a non-zero exit from the CLI
find_similar(…, exclude=…) exclude into ONNX Runtime's session options a call keyword, like run() already treated it
read_labels('model.onnx') 100,877 class names of binary refuses anything but .txt, .csv, .json

That --conf one: tools.py has from __future__ import annotations, so every annotation reaches _coerce as a string and no CLI flag was ever converted. The old test passed the real int and so tested nothing the CLI does.

Evidence that the answers are right, not just different

Nothing here asserts against numbers written down by hand. Two detector families check each other, and the classifier checks itself against the transform its weights were evaluated with.

  • anya's tensor for mobilenetv4_conv_small equals timm's torchvision eval transform to max|diff| 0.000000, and SegFormer's and DINOv2's own processors to 0.0175, which is one uint8 step over the ImageNet std.
  • A quantised efficientdet and a float yolov8n, different architectures on different runtimes, now put the same boxes in the same places: the bus in bus.jpg to max|diff| 31px of an 800px-wide box, and the two cats in 000000039769.jpg to 13.5px and 11.7px. Before the SSD padding fix the nearer cat was 47px out in y.
  • A 440Hz sine through yamnet is Sine wave at 0.996; silence is Silence; noise is Static.
  • lraspp on VOC gives class 8, cat, 36.8% of a picture of two cats. SegFormer on ADE20K gives building, sidewalk, bus for a photo of a bus.
13/13 passed

The eval is re-runnable and clean-runnable, which took fixing twice: once because classify over a folder read the sorted tree a previous run had written into it, and once because find_similar did.

Structure

  • Model._finish resolves the labels, the task and the Prep for every runtime. All three repeated the same eight lines, and each new prep keyword had to be threaded through three files. _own_labels, _guess_task and _mk_prep are the hooks; ONNX overrides none of them. This is what CLAUDE.md already said the architecture was.
  • read_labels takes a dict and inverts label2id, so hub.config_labels is gone.
  • with_hub_defaults returns prep keywords rather than a fixed 4-tuple, so crop_pct and resample needed no new slot anywhere.
  • find_models searches each word after the phrase, in search order: the Hub matches search against the repo id, so bird classifier with runtime='litert' found nothing where bird found a repo, and sorting the merged results by downloads buried it again under repos matching classifier.

The environment

.cursor/environment.json runs .cursor/install.sh, which installs the dev group rather than restating it: eleven seconds on a fresh pod, two warm. Two draft builds off this branch both succeeded, and two fresh agents booted from them ran nbdev-export clean, nbdev-test green, evals/e2e.py twice at 13/13, the CLI, and the install script again for idempotence.

The eval takes about six seconds on a pod booted from the tested build rather than downloading 190MB, because the snapshot carries the warmed HuggingFace cache. It downloads unauthenticated; an HF_TOKEN secret is not needed but would remove any rate-limit risk on a busy pod.

Each runtime notebook and the index now show a real repo id, marked eval: false, with the answers the eval asserts on. nbdev-test is green and nbdev-export is clean. No fixture changed shape.

Open in Web Open in Cursor 

cursoragent and others added 15 commits August 13, 2026 21:12
Every axis of an optimum-exported ONNX input is symbolic and named, so the channel
axis has to be read from the names rather than the sizes: three of the four repos
tested fed NHWC pixels into an NCHW graph. A size the graph states outright now
wins over one a preprocessor_config asks for, since the graph will take no other.
'logits' names the output of a classifier and a segmenter alike and so decides
nothing; a 'last_hidden_state' is features. SegFormer's 150 classes on a 128x128
grid are channels-first, and its label map goes back onto the original picture.
An embedding output that is a token sequence or a feature map is mean-pooled to
one vector instead of flattened. crop_size with size means timm's crop_pct.
with_hub_defaults returns prep keywords, so a new one needs no new tuple slot.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Every repo tested asks for bicubic or bilinear in `resample`, and anya used bilinear
for both. On five photos through mobilenetv4 the filter flipped top-1 once, at
0.336 against 0.340, so the config's choice is carried through to PIL. With it,
anya's tensor for mobilenetv4 matches timm's torchvision eval transform exactly
(max|diff| 0.00000) and SegFormer's and DINOv2's processors to 0.0175, which is
one uint8 step divided by the ImageNet std.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Boxes: the same (1, 84, 8400) signature carries pixels in one export and 0..1 in
the next, so the unit is read off the range instead of assumed. SSD boxes are
normalised against the padded input, so they now go through scale_boxes like every
other detector's; efficientdet's cat box on COCO 39769 moves from [5,99,315,400] to
[5,52,315,453], against yolov8n's [12,56,319,466] for the same cat.

ssdlite320's tflite ships twelve raw per-stride heads, which the three-or-more rule
read as a postprocess head and silently decoded to nothing. is_ssd checks the
boxes/classes/scores contract and anything else raises.

Labels: a labels.txt sits at the repo root while the weights sit under onnx/ or
tflite/, so sidecar_labels walks up, and both runtimes use it. read_labels takes a
config dict and inverts label2id, which is what config_labels did, so that is gone.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
tools.py has `from __future__ import annotations`, so every annotation reaches
_coerce as the string 'int' rather than int, and no flag was ever converted:
--topk=2 reached np.argsort as '2' and --conf=0.6 reached a numpy comparison as
'0.6'. The old test passed the real type and so tested nothing the CLI does.

find_models searches each word after the phrase, because the Hub matches search
against the repo id and 'bird classifier' with runtime='litert' found nothing
where 'bird' found a repo. Search order is kept over download count, or the
20k-download 'classifier' matches bury the one that matched 'bird'.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
…it is called

is_prob required the scores to sum to 1, so yamnet's 521 sigmoid scores, which sum
to 4, were softmaxed a second time: silence read as Silence 0.004 and now reads
0.801, noise as Static 0.003 and now 0.500. Being in 0..1 is the test, and since
softmax is monotonic this only ever changes the reported score, never the ranking.

decode_classify averages a per-frame output, which is what yamnet emits every
0.48s of audio. Three output tensors are a detector only when the first is four
columns wide, or yamnet's scores/embeddings/spectrogram read as one.

Label files are named labels_yamnet.txt and audioset_labels.txt as often as
labels.txt, and sit at the repo root while the weights sit under onnx/. Matching
the shape of the name rather than a fixed list gives EdgeFirst/yolo11-det its 80
COCO names, where it reported class_15 for a bus before. fetch_sidecars took a
dest it never wrote to.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
evals/e2e.py downloads eight repos and asserts on the answers: a timm classifier
whose axes are all symbolic, a torch tflite classifier with no config at all, a
float16 yolov8n in pixels, a tflite yolov8 in 0..1, a quantised efficientdet with
a postprocess head, SegFormer on ADE20K, lraspp on VOC, DINOv2, and yamnet on a
sine. Two detector families are checked against each other's boxes for the same
object rather than against numbers written down here, and the classifier's tensor
against timm's own torchvision transform, which it matches to 1e-6.

Fixes it turned up: the runtimes did not accept crop_pct or resample, so both went
into **kw and on to the Interpreter constructor. infer_task called ssdlite320's
twelve heads a classifier once yamnet taught it that three heads can be one; it
now raises at load, naming task=, rather than decoding nonsense at predict.
read_labels read whatever file it was handed, so a path to model.onnx became
100877 class names. TOOLS are wrapped so a chat gets {error} and not a traceback.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
.cursor/install.sh builds .venv with all four inference extras, nbdev, and the CPU
torch and transformers wheels that evals/e2e.py compares anya's preprocessing
against. Two seconds warm, and rerunnable. nbdev's git hooks are installed, since
the source here is notebooks.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
… of a subclass

The notebooks only ever met the toy graphs, so a reader saw the claim that a hub
repo id changes nothing and no evidence of it. Each runtime notebook and the index
now show one, marked eval: false, with the answers evals/e2e.py asserts on.

The Model section said subclasses supply _load and spec. They supply _read_spec,
_infer and a Prep, which is what CLAUDE.md says and what the code does.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
All three runtimes repeated the same eight lines after reading their signature,
and each new prep keyword had to be threaded through three files. Model._finish
does it once; _own_labels, _guess_task and _mk_prep are the hooks, and ONNX now
overrides none of them. Core ML reads its signature in a _read_spec like the
others and picks up a sidecar labels file, which it never did.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
run() splits **kw into call keywords and model keywords; these two did not, so
find_similar(q, folder, model, exclude=sorted) put exclude into the Model
constructor and on into the runtime's session options. split_kw does it once for
all three. Found by the eval: find_similar over a folder that had been sorted in
place answered 1.0 five times, once for each hard link.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
pyproject already lists the runtimes and nbdev under [dependency-groups] dev, so
the install script listing them again could drift from what a contributor gets.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
The tools return {error} rather than raising, so the exit code was the only thing
left telling a shell that anya detect_image found no model.

Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants