Repository navigation
Run anya against real models, and fix the twelve things that broke - #1
Draft
Karthik777 wants to merge 15 commits into
Draft
Karthik777 wants to merge 15 commits into
Karthik777 wants to merge 15 commits into
Conversation
Every axis of an optimum-exported ONNX input is symbolic and named, so the channel axis has to be read from the names rather than the sizes: three of the four repos tested fed NHWC pixels into an NCHW graph. A size the graph states outright now wins over one a preprocessor_config asks for, since the graph will take no other. 'logits' names the output of a classifier and a segmenter alike and so decides nothing; a 'last_hidden_state' is features. SegFormer's 150 classes on a 128x128 grid are channels-first, and its label map goes back onto the original picture. An embedding output that is a token sequence or a feature map is mean-pooled to one vector instead of flattened. crop_size with size means timm's crop_pct. with_hub_defaults returns prep keywords, so a new one needs no new tuple slot. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Every repo tested asks for bicubic or bilinear in `resample`, and anya used bilinear for both. On five photos through mobilenetv4 the filter flipped top-1 once, at 0.336 against 0.340, so the config's choice is carried through to PIL. With it, anya's tensor for mobilenetv4 matches timm's torchvision eval transform exactly (max|diff| 0.00000) and SegFormer's and DINOv2's processors to 0.0175, which is one uint8 step divided by the ImageNet std. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Boxes: the same (1, 84, 8400) signature carries pixels in one export and 0..1 in the next, so the unit is read off the range instead of assumed. SSD boxes are normalised against the padded input, so they now go through scale_boxes like every other detector's; efficientdet's cat box on COCO 39769 moves from [5,99,315,400] to [5,52,315,453], against yolov8n's [12,56,319,466] for the same cat. ssdlite320's tflite ships twelve raw per-stride heads, which the three-or-more rule read as a postprocess head and silently decoded to nothing. is_ssd checks the boxes/classes/scores contract and anything else raises. Labels: a labels.txt sits at the repo root while the weights sit under onnx/ or tflite/, so sidecar_labels walks up, and both runtimes use it. read_labels takes a config dict and inverts label2id, which is what config_labels did, so that is gone. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
tools.py has `from __future__ import annotations`, so every annotation reaches _coerce as the string 'int' rather than int, and no flag was ever converted: --topk=2 reached np.argsort as '2' and --conf=0.6 reached a numpy comparison as '0.6'. The old test passed the real type and so tested nothing the CLI does. find_models searches each word after the phrase, because the Hub matches search against the repo id and 'bird classifier' with runtime='litert' found nothing where 'bird' found a repo. Search order is kept over download count, or the 20k-download 'classifier' matches bury the one that matched 'bird'. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
…it is called is_prob required the scores to sum to 1, so yamnet's 521 sigmoid scores, which sum to 4, were softmaxed a second time: silence read as Silence 0.004 and now reads 0.801, noise as Static 0.003 and now 0.500. Being in 0..1 is the test, and since softmax is monotonic this only ever changes the reported score, never the ranking. decode_classify averages a per-frame output, which is what yamnet emits every 0.48s of audio. Three output tensors are a detector only when the first is four columns wide, or yamnet's scores/embeddings/spectrogram read as one. Label files are named labels_yamnet.txt and audioset_labels.txt as often as labels.txt, and sit at the repo root while the weights sit under onnx/. Matching the shape of the name rather than a fixed list gives EdgeFirst/yolo11-det its 80 COCO names, where it reported class_15 for a bus before. fetch_sidecars took a dest it never wrote to. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
evals/e2e.py downloads eight repos and asserts on the answers: a timm classifier
whose axes are all symbolic, a torch tflite classifier with no config at all, a
float16 yolov8n in pixels, a tflite yolov8 in 0..1, a quantised efficientdet with
a postprocess head, SegFormer on ADE20K, lraspp on VOC, DINOv2, and yamnet on a
sine. Two detector families are checked against each other's boxes for the same
object rather than against numbers written down here, and the classifier's tensor
against timm's own torchvision transform, which it matches to 1e-6.
Fixes it turned up: the runtimes did not accept crop_pct or resample, so both went
into **kw and on to the Interpreter constructor. infer_task called ssdlite320's
twelve heads a classifier once yamnet taught it that three heads can be one; it
now raises at load, naming task=, rather than decoding nonsense at predict.
read_labels read whatever file it was handed, so a path to model.onnx became
100877 class names. TOOLS are wrapped so a chat gets {error} and not a traceback.
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
.cursor/install.sh builds .venv with all four inference extras, nbdev, and the CPU torch and transformers wheels that evals/e2e.py compares anya's preprocessing against. Two seconds warm, and rerunnable. nbdev's git hooks are installed, since the source here is notebooks. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
… of a subclass The notebooks only ever met the toy graphs, so a reader saw the claim that a hub repo id changes nothing and no evidence of it. Each runtime notebook and the index now show one, marked eval: false, with the answers evals/e2e.py asserts on. The Model section said subclasses supply _load and spec. They supply _read_spec, _infer and a Prep, which is what CLAUDE.md says and what the code does. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
All three runtimes repeated the same eight lines after reading their signature, and each new prep keyword had to be threaded through three files. Model._finish does it once; _own_labels, _guess_task and _mk_prep are the hooks, and ONNX now overrides none of them. Core ML reads its signature in a _read_spec like the others and picks up a sidecar labels file, which it never did. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
run() splits **kw into call keywords and model keywords; these two did not, so find_similar(q, folder, model, exclude=sorted) put exclude into the Model constructor and on into the runtime's session options. split_kw does it once for all three. Found by the eval: find_similar over a folder that had been sorted in place answered 1.0 five times, once for each hard link. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
pyproject already lists the runtimes and nbdev under [dependency-groups] dev, so the install script listing them again could drift from what a contributor gets. Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
The tools return {error} rather than raising, so the exit code was the only thing
left telling a shell that anya detect_image found no model.
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
Co-authored-by: Karthik Rajgopal <Karthik777@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The tiny graphs in
nbs/fixturesprove the plumbing. They say nothing about what a real export declares, and it turns out most real exports declare something anya read wrong.evals/e2e.pydownloads eight repos, runs them over five COCO and Ultralytics photographs plus three generated sounds, and asserts on the answers. Everything below was found by that script failing.The one that mattered most
Three of the four ONNX repos tested declare their input as
['batch_size', 'num_channels', 'height', 'width']— every axis symbolic, and named. anya read the channel axis from the sizes, found no integers, guessed NHWC, and fed channels-last pixels into a channels-first graph:chan_axisnow reads the names when the sizes cannot say. This is the commonest shape optimum emits, so it is the difference between anya working on the Hub and not.The rest
webnn/yolov8nsizepreprocessor_config.json, into a graph fixed at 640Xenova/segformer-…taskclassify, because the output is namedlogitssegment;logitsis what both call their output and so decides nothingXenova/dinov2-smallcrop_size224 withsize256crop_pctresamplescale_boxes, like every other detectorssdlite320's twelve raw heads[]task=labels.txtat the repo root, weights underonnx/class_15for a bus--conf=0.6'0.6'into a numpy comparison{"error": …}, and a non-zero exit from the CLIfind_similar(…, exclude=…)excludeinto ONNX Runtime's session optionsrun()already treated itread_labels('model.onnx').txt,.csv,.jsonThat
--confone:tools.pyhasfrom __future__ import annotations, so every annotation reaches_coerceas a string and no CLI flag was ever converted. The old test passed the realintand so tested nothing the CLI does.Evidence that the answers are right, not just different
Nothing here asserts against numbers written down by hand. Two detector families check each other, and the classifier checks itself against the transform its weights were evaluated with.
mobilenetv4_conv_smallequals timm's torchvision eval transform tomax|diff| 0.000000, and SegFormer's and DINOv2's own processors to 0.0175, which is one uint8 step over the ImageNet std.bus.jpgtomax|diff| 31pxof an 800px-wide box, and the two cats in000000039769.jpgto 13.5px and 11.7px. Before the SSD padding fix the nearer cat was 47px out in y.Sine waveat 0.996; silence isSilence; noise isStatic.lrasppon VOC gives class 8, cat, 36.8% of a picture of two cats. SegFormer on ADE20K gives building, sidewalk, bus for a photo of a bus.The eval is re-runnable and clean-runnable, which took fixing twice: once because
classifyover a folder read the sorted tree a previous run had written into it, and once becausefind_similardid.Structure
Model._finishresolves the labels, the task and thePrepfor every runtime. All three repeated the same eight lines, and each new prep keyword had to be threaded through three files._own_labels,_guess_taskand_mk_prepare the hooks; ONNX overrides none of them. This is what CLAUDE.md already said the architecture was.read_labelstakes a dict and invertslabel2id, sohub.config_labelsis gone.with_hub_defaultsreturns prep keywords rather than a fixed 4-tuple, socrop_pctandresampleneeded no new slot anywhere.find_modelssearches each word after the phrase, in search order: the Hub matchessearchagainst the repo id, sobird classifierwithruntime='litert'found nothing wherebirdfound a repo, and sorting the merged results by downloads buried it again under repos matchingclassifier.The environment
.cursor/environment.jsonruns.cursor/install.sh, which installs thedevgroup rather than restating it: eleven seconds on a fresh pod, two warm. Two draft builds off this branch both succeeded, and two fresh agents booted from them rannbdev-exportclean,nbdev-testgreen,evals/e2e.pytwice at 13/13, the CLI, and the install script again for idempotence.The eval takes about six seconds on a pod booted from the tested build rather than downloading 190MB, because the snapshot carries the warmed HuggingFace cache. It downloads unauthenticated; an
HF_TOKENsecret is not needed but would remove any rate-limit risk on a busy pod.Each runtime notebook and the index now show a real repo id, marked
eval: false, with the answers the eval asserts on.nbdev-testis green andnbdev-exportis clean. No fixture changed shape.