Repository navigation
[RNE Rewrite] feat: benchmark harness for every registry variant - #1363
msluszniak wants to merge 32 commits into
Conversation
383f54e to
df61e4d
Compare
Running this on your devicesgit submodule update --init --recursive # phonemis, CMake fails without it
yarn install
cd apps/benchmarks
yarn bench --platform android --suite quick --label v0.10.0 --max-temp-c 37
yarn bench --platform ios --suite quick --label v0.10.0
yarn bench:summary <out>.jsonlSend back the Notes:
|
df61e4d to
9dc8c16
Compare
|
Pushed a0fb58b: the The profile was zeroed before the warmups and divided by Fix: zero the profile between the warmups and the timed loop, inside the worklet.
|
bench — Pixel 10
bench — iPhone SE (3rd gen)
bench — iPhone 17
Rows appended by @msluszniak from the raw |
|
Thank you @barhanc, during the run for llms, I got this error for vulkan model: Which is probably because vulkan does not correctly handle |
v0.10.0 on SM-S948B (Android 16, SM8850, 8 cores, 11 GB)The published Android estate: all 124 variants
classification
style-transfer
semantic-segmentation
keypoint-detection
object-detection
instance-segmentationPost-processing is mask-bound; read Execute %.
privacy-filter
text-embeddings
image-embeddings
ocrRecognizer cost scales with the regions the detector finds.
voice-activity-detection
speech-to-textDecode length is input-dependent; read Execute %.
text-to-speechTimed on the RN thread; includes per-chunk thread hops. The nine Kokoro entries are three networks, not nine: Read the median for
text-to-image
LLMs: 23 measured, 1 failed, 15 not runNot comparable to the table above. Decode is pinned to 64 tokens, EOS ignored, temperature 0, cold prompt each iteration, 10 timed iterations after 2 warmups (not 20/3) since one iteration is a whole generation.
* Vulkan LLMs fail, and the reported error was a decoy. Investigated, root-caused, fix in #1454. The reported The real failure, once unmasked: 262144 is 256x1024 and 274432 is 268x1024: the prefill input is bounded at 256 tokens. The Vulkan LLM exports are dynamic-shape, where
All three Vulkan LLMs are affected, Gemma4 worst. Only Vulkan is worth having once that lands: in isolation the same model ran 505 ms against XNNPACK's 1692 ms for 64 tokens, 125 against 37 tok/s, a 3.3x win. The medians above still measure a growing prompt. Harness bugs. Failures used to be written twice, since the collector's duplicate guard only covered |
|
Found a defect in the thermal gate while running an iOS sweep today, plus three smaller fixes. All are one-file changes against this branch. iOS runs sleep 90s per iteration regardless of temperature
const BLIND_SETTLE_MS = 90_000;
...
if (state.status <= 0 && coolEnough) {
if (!readable) await sleep(BLIND_SETTLE_MS);
Measured cost on one model, The collector is what triggers it: Three other fixes from the same runA configured sink that never answers should fail, not warn.
Happy to open these as a PR against this branch if you prefer that to a patch dump. |
2ca40f9 to
48ca5b3
Compare
v0.10.0 on Samsung Galaxy S20+ (Android 13, SM-G986B, 8 cores, 11 GB, Exynos 990, Mali-G77 MP11)The published Android estate: 124 variants
classification
style-transfer
semantic-segmentation
keypoint-detection
object-detection
instance-segmentation
privacy-filter
text-embeddings
image-embeddings
ocr
voice-activity-detection
speech-to-textDecode length is input-dependent; read Execute %.
text-to-speech
text-to-image
|
v0.10.0 on iPhone 17 (iOS 26.6.2, iPhone18,3, 6 cores, 8 GB)The published iOS estate: 190 variants
classification
style-transfer
semantic-segmentation
keypoint-detection
object-detection
instance-segmentation
privacy-filter
text-embeddings
image-embeddings
ocr
voice-activity-detection
speech-to-textDecode length is input-dependent; read Execute %.
text-to-speech
text-to-image
Failed (1)
Common causes:
Error: Error::InvalidProgram |
|
@barhanc regarding this one:
I re-exported this model and how the one on HF should be the correct one. You can re-try it. |
pixel10-v1 on Pixel 10 (Android, Tensor G5, 8 cores, GB)The published estate: 124 variants Thermal gate: 40°C ceiling (raised from 37°C because the device was charging and idled at 38.5–39°C). Measurements were taken while the device was USB-connected and charging, so sustained throughput may differ from a battery-powered run. Thermal throttling checks were disabled to avoid indefinite gate stalls under charging.
classification
style-transfer
semantic-segmentation
keypoint-detection
object-detection
instance-segmentation
privacy-filter
text-embeddings
image-embeddings
ocr
voice-activity-detection
speech-to-textDecode length is input-dependent; read Execute %.
text-to-speech
text-to-image
|
|
Thank you @barhanc. Do we have a complete list now? |
iphone-se-v1 on iPhone14,6 (ios, , 6 cores, GB)The published estate: 187 variants
classification
style-transfer
semantic-segmentation
keypoint-detection
object-detection
instance-segmentation
privacy-filter
text-embeddings
image-embeddings
ocr
voice-activity-detection
speech-to-textDecode length is input-dependent; read Execute %.
text-to-speech
Failed (7)
|
|
I will benchmark LLMs after runner refactor. |
The biggest source of noise on a phone is the clock, not the code: a device boosts early and sags as it heats, which is how two runs of identical code came out 9% to 51% apart. Android exposes PowerManager's fixed-performance mode over `cmd power`, and vendors implement it as a hard frequency cap rather than a hint. On the S26 Ultra it takes every cluster from 3.19/3.40 GHz to about 1.98 GHz. Absolute numbers drop, which is the trade: the same clock in every run is worth more than a fast one. `--pin-clocks` defaults to `auto` (pin where supported), with `on` to require it and `off` to opt out. The driver reads the frequency back rather than trusting the call, since not every vendor implements the HAL, and restores normal clocks on exit, signal and crash alike so a capped device is never left behind. The run records whether it was pinned and the comparator refuses to diff a pinned run against an unpinned one. No iOS equivalent exists; nothing in the public API pins the clock, so runs there still depend on the thermal gate.
A fixed `--cooldown 420` is wrong in both directions: it burns seven minutes on a phone that is already cold, and it is not enough after a heavy suite. `--cooldown auto` polls `dumpsys` until the framework reports no throttling and the battery temperature has stopped falling. It waits on a plateau rather than an absolute threshold, because what counts as cool differs per device while "no longer dropping" does not, and it requires two consecutive settled samples so a flat reading mid-fall does not end the wait early. A 30s floor lets the heat of building and installing dissipate; `--cooldown-max` stops a warm room or a charging phone stalling the run forever. Charging is reported, since it keeps a device warm. On an idle S26 Ultra this releases after 30s where the fixed wait took 420s. Android only: iOS exposes no thermal readout to the host, so `auto` falls back to a fixed sleep there rather than pretending to measure.
The FastSAM case wedged a run. FastSAM pairs a 0.5 confidence threshold with an IoU of 0.9, and NMS at 0.9 suppresses almost nothing, so on a textured synthetic image nearly every candidate survives and each survivor materialises a full 640x640 mask in JS. The worklet thread stalled with no output. RF-DETR nano emits a fixed set of queries and runs NMS at 0.55, so its post-processing is bounded whatever the input looks like. Verified on an S26 Ultra: 215 ms median, 516 MB peak. This is the input-dependence the README already warns about, met head on: a model whose post-processing cost is unbounded in the number of detections does not belong in a suite fed deliberately adversarial synthetic images.
The suite measured 17 hand-written cases against a registry that publishes 261 variants. Extending it by hand does not scale and does not stay correct: a variant added to models.ts and not to the suite is a model that silently never gets benchmarked, which is the failure this harness exists to prevent. So the case list is derived rather than written. generate-variants.mjs evaluates models.ts under Node's type stripping and emits every concrete variant with its backend, precision, platforms and download size; suite.ts joins each to a per-task driver. 163 variants are runnable on Android and 234 on iOS. Adding a model or a variant now needs no change here at all; adding a task needs one driver; a task with no driver is reported as skipped rather than dropped. Each variant is measured three times, and each measurement starts with the device at or below 35C. The gate is an absolute ceiling rather than the plateau rule it replaces: a plateau answers "has it stopped cooling", which is the right question for two runs on one device and the wrong one for four devices, since a phone settling at 41C and one settling at 30C both pass it. The wait lives on the host because Android exposes battery temperature to adb and not to an app; iOS has no readout at all, so it falls back to thermalState plus a fixed settle and records that it did, rather than implying 35C. Repeats are not iterations. Iterations bound the noise inside one measurement; repeats expose the run-to-run spread that thermal state and clock drift produce, which on a phone is the larger of the two. The comparator folds repeats to a median and widens each metric's noise floor by the across-repeat range, so a thermal artefact stops reading as a regression. Measurements are appended to a JSONL as they land and --resume skips what it already holds: a whole-estate run is hours, and a report assembled only at the end loses all of it to a crash on the last case. Models are deleted after a variant's last repeat, so peak disk is one model rather than 119 GB, and a variant over --max-bytes is recorded as skipped with its size instead of failing halfway through its download. Also adds an LLM driver (decode pinned to 64 tokens with EOS ignored, since generation length is a property of the model and not of the runtime), a summarize script for the table people actually read, and BENCHMARK_SPEC.md as the protocol other devices run against.
Fixed-performance mode caps a Galaxy S26 Ultra from 3.19/3.40 GHz to about 1.98 GHz. That is the right trade for detecting a regression, where the same clock in both runs is worth more than a fast one, and the wrong one for publishing device numbers: every figure would understate the phone by roughly the ratio of the clocks and describe a state its governor would never choose. The flag stays, for A/B work against another build. What changes is which way it points when nobody says. With the clock free the thermal gate is the only control left over run-to-run drift, which is what the per-repeat gate is for. Also resolves a contradiction in the spec, which told people to charge the phone for a long run two sections after telling them charging stalls the gate.
Repeating a whole measurement was meant to expose the run-to-run spread that thermal state and clock drift produce, on the assumption it was larger than the spread inside one measurement. Measured on a Galaxy S26 Ultra it is not: 16.8% within against 12.1% across on EfficientNet int8, 21.7% against 19.6% on fp32. Twenty back-to-back iterations already show what three cold repeats show, at roughly a third of the wall clock, because each repeat waits at the gate again. The spread that remains belongs to the metric rather than to the sampling. pipeline carries garbage collection in its TypeScript post-processing and sits at 7-22%; the raw execute figure is ExecuTorch alone and sits near 2.5%. No repeat count changes that, which is why the summary now leads with where the time goes rather than with an error bar. --repeats 3 still does the old thing for a model worth pricing precisely. Also reports Model MB (peak minus the baseline taken just before the load) and Execute %, after finding that absolute peak charges a model measured late for the cases before it: every case shares one process, and its baseline crept from 292 MB to 350 MB over seven measurements.
Two ways a run could quietly measure under settings nobody chose. expo run:* attaches to a bundler already listening rather than starting one, and every EXPO_PUBLIC_BENCH_* value is inlined at transform time, so the settings live in that process. A run asked for one repeat under a new label and the app announced three repeats under the previous one. The driver now frees the dev-server port first, and the app echoes the label and repeat count it actually booted with so a mismatch is a hard stop rather than a footnote: a complete set of plausible numbers taken under the wrong configuration is worse than a crash. The device-side gate then ignored the ceiling it was given. It exists for iOS, where nothing exposes a temperature, but it also catches any blip in the adb tunnel — and on Android BenchProbe does report one. A lost /gate POST dropped a measurement onto that path and it started at 36.6C under a 35C gate, with the reading sitting in its own result. It now holds the ceiling wherever a temperature is readable, and keeps the blind settle only where one is not.
A ceiling below the device's idle floor never opens. A Galaxy S26 Ultra sits at 35.4C doing nothing with the harness in the foreground, because the screen is held on for the length of the run or Android freezes the app mid-suite. Under a 35C gate every measurement waited out its full 30 minute timeout and then measured warm regardless, which would have turned a 52 model tier into a day of waiting and made the timedOut flag meaningless by setting it on every row. 37C opens immediately on an idle device and still holds after a model that heated it. It stays an absolute number, since two phones gated differently are not comparable, which is the whole reason for a fixed ceiling over a plateau rule. --resume was also weaker than it read. The app asks the collector which measurements exist and skips them, over the same adb tunnel that has already been seen to drop a POST: a lost answer means everything is measured again and the file quietly grows a second copy of every row. The collector now holds the keys its output file already contains and refuses one it has, so the guarantee belongs to the writer rather than to a request surviving.
Stopping the charge was meant to keep the device under a 35C gate, since the cable adb needs holds it about 1.5C warmer. It worked, and it invalidated the measurements: dumpsys battery unplug convinces the framework the device is on battery, Samsung answers by engaging Battery Saver, and it sticks until the cell reaches 90%. Nothing about it looks like throttling — the CPU's maximum frequencies read normal — but EfficientNet int8 went from 66 ms to 116 ms with nothing else changed. A phone in a power-saving mode no user asked for is not the phone the numbers describe. With the gate at 37C the trade is unnecessary: charging costs 1.5C and the ceiling has room for it, so runs stay plugged in and the battery survives a long suite. --unplug is still there for a device whose idle floor needs it, and now clears low_power rather than leaving it set. The run also refuses to start in Battery Saver rather than trusting that nobody turned it on, because the harness itself turned it on once and the resulting numbers looked entirely plausible.
The suite answers "how fast". This answers "what did the model actually compute on this device", which is a question a host cannot answer for #1406: Core ML fp16 draws shifted masks and boxes on an iPhone's ANE, and a Mac cannot reproduce it because its own ANE compiler rejects the model and Core ML falls back to CPU/GPU without saying so. Every host check therefore passed. The host writes the input as a raw tensor and serves it; the device feeds those bytes to execute untouched and sends the output tensors back as base64 of their own buffers. Neither side decodes an image or rounds through decimal text, so a difference in the output cannot be a difference in preprocessing or in serialisation - which matters when the signal being measured is a fraction of a pixel. iOS needs the ATS exception and a local-network usage string to reach a collector on the LAN at all; Android already had usesCleartextTraffic.
Two defects made the execute share unusable, and both are fixed here. The share was derived from a standalone replay that sized every tensor at its schema maximum. For a model with a dynamic dimension that is not the work the pipeline did: a text embedder declaring 510 tokens was replayed at 510 while the pipeline ran the ~75 of the benchmark text, so execute exceeded the pipeline and the column was blanked for 21 of 52 variants. OCR was worse, since the pipeline calls `recognize` once per detected box and the replay called it once. `getExecutionProfile()` accumulates time inside `Model::execute` instead, so the figure covers exactly the shapes and call counts the pipeline used, and the share is always defined. The API is new public surface because a task pipeline owns its `Model` privately and callers had no way to ask. The runs were also debug builds. ExecuTorch is a prebuilt release library so `execute` was unaffected, but the library's own C++ compiles unoptimised and JS is served as a dev bundle, inflating everything around the model by roughly an order of magnitude. That does not merely add noise, it inverts the conclusion: EfficientNet measured 29% ExecuTorch in debug and 91% in release. Release is now the default, debug warns, and the build type is recorded in the report. Drops the ET 1.3.1 baseline: schema 1, debug, no thermal gate, and no execution data, so nothing the harness produces can be compared to it.
The spec is the contract other devices follow, and it did not say which build type to use. A debug run does not merely read slow, it inverts the verdict: EfficientNet is 29% ExecuTorch in debug and 91% in release.
…pass Cuts about 4,000 lines from the harness without losing a measurement. `src/variants.generated.ts` was 3,171 checked-in lines produced by a script, plus a CI check to catch it going stale. The registry is a plain nested object once imported and `variants()` only spreads a map, so the list is a walk, not something that needs generating. Everything in the generated file was derivable in-process except the download sizes, which need a network round trip, so those alone stay cached and the generator shrinks to the script that measures them. The replay pass goes with it. It loaded each `.pte` separately and sized tensors at the schema maximum, which is not what the pipeline ran: it is what made the execute share unusable for 21 of 52 variants and produced a 510-token forward to divide by a 75-token pipeline. The in-band profiler measures the real work, so the replay answered no question worth the code. The RF-DETR ANE probe moves out to its own branch. It rides on the collector but it is a device-only correctness investigation for #1406, not part of a benchmark harness.
Both functions wrap a native JSI call, so the repo's convention requires the directive or they cannot be serialized onto a worklet runtime. That is not hypothetical here: the profile is read around a measurement loop that runs inside one, which is also why the C++ side takes a mutex. Records the three new exports in the API surface snapshot.
The execution profile was zeroed before the warmups and then divided by iterations + warmup, so "Execute ms" was a mean over all 23 passes while "Inference ms" was the median of the 20 timed ones. A warmup pays one-off costs the timed median never sees, Vulkan shader compilation above all, so ExecuTorch time came out larger than the pipeline time that contains it on 23 of 52 variants - MiniLM on Vulkan fp16 by 1.47x, which works out to roughly 18 ms per warmup pass. The profile is now zeroed between the warmups and the timed loop, inside the worklet, so the tally covers exactly the iterations the durations cover. That works because resetExecutionProfile carries the 'worklet' directive. summarize.mjs compared that per-iteration mean against the pipeline median, mixing two statistics; it now compares mean to mean. It also clamped the share to 100% and floored JS ms at 0, which is how a physically impossible reading rendered as a tidy "100% / 0.00 ms" and went unnoticed. Both clamps are gone: a share over 100% is a measurement fault and has to be visible. Verified on the first re-run rows: mosaic-int8 101% -> 99%, and every row now sits under 100%.
…e case The gate ran once per case, which fixed the temperature a measurement started at and said nothing about the one it ran at. A case whose iterations take seconds heats the phone as it goes: `smollm2-1.7b` entered its window at 36.6 C, reached `severe` at 41.5 C partway through, and returned iterations spanning 2.6 s to 7.3 s. The end-of-case thermal reading had recovered by then, so nothing in the row said the number described a throttling phone. Iterations on the async path are now held at the gate before each one, warmups included, so the whole pass stays inside the ceiling instead of drifting out of it. The wait sits outside the timed window: it is the harness's cost, not the model's. Cases on the worklet path still gate per case, since holding inside the worklet loop would add a thread hop per iteration and change what is being measured. What the device did during the pass is now recorded too, as the worst state seen rather than the one left at the end, and `thermalValid` says whether the holds kept it there. Also here, all found while measuring the LLM tier: - `--order size` runs the smallest download first, so an interrupted sweep has covered the cheap end rather than nothing. - The collector wrote a failure twice, because its duplicate guard only tracked successes. - `--resume` resolves its file from `--label` and `--out`, so changing either silently re-measured the whole suite. It now says which file it read and that it found nothing.
`thermalValid` asked whether the device stayed under the ceiling for the whole pass, which no LLM row can satisfy: a second of all-core decode heats this phone several degrees, so an iteration crosses the ceiling while it runs however cold it began. The first re-measurement under per-iteration holds entered at 32.9 C and still peaked at 38.9 C, so the flag condemned a measurement that was sound. Where an iteration ends is not in the harness's gift; where it starts is. The flag now reports whether every hold reached the ceiling rather than giving up, and `holdsTimedOut` counts the ones that did. How hot the pass actually got is still recorded in `thermalPeak`, for a reader to weigh. Worth keeping in view: the holds are doing real work. The same model measured 4593 ms while throttling and 3624 ms held, with the spread halved.
expo run:* starts a bundler even for a release build, and it was always the default 8081. A second concurrent run therefore either attached to the first run's bundler, inheriting the EXPO_PUBLIC_BENCH_* values it was launched with, or killed it: resetBundler kill -9'd whatever held the port. --dev-port makes that port per-run, and parseArgs rejects it colliding with the collector port rather than failing partway through a build. No fan-out mode: the reason to run two devices together is usually that they need different case lists, which --only/--tasks/--suite already express per invocation.
…collector EXPO_PUBLIC_BENCH_URL_MAP swaps a published model URL for a local one before download, so a local export is fetched, cached and loaded on exactly the path a registry file takes. That is what makes an A/B against a published artifact honest: only the bytes differ. A configured sink that has never accepted a POST now throws instead of warning. It was warning once per measurement while discarding every result, so a run could complete having saved nothing. A sink that worked and later dropped still degrades to a warning, since the final report re-sends everything.
The raw execute pass is gone, so drivers no longer need modelPathKey. Also removes nativeHeapBytes, freeDiskBytes, ranWithoutThrottling and the library's totalExecutionMs export, none of which had a caller. The thermal watch now treats iOS's -1 sentinels as unknown instead of as readings.
77 published files had no cached size after the rebase. The script also ran main() twice and kept sizes for URLs the registry no longer publishes. Aligns react-native-blob-util with the other apps and keeps the wordlist append-only.
- The collector writes a run header (device, build type, clock pinning, settings, URL map) into the JSONL, so compare and summarize read it directly. Resume refuses a file recorded under different settings, and a fresh run refuses to write into a file that already holds measurements. - On Android a dead app process is detected, the case it died on is recorded as an error, and the app is relaunched to continue. - The final JSON is rebuilt from the JSONL, so a resumed run's report is complete. - compare: per-method execute rows, call counts as a workload check, execute folded across repeats, warm cases withheld per case instead of refusing the whole comparison, build type mismatch is fatal. - iOS was never gated: the host answered every gate request immediately. It now waits for a nominal thermal state. - --url-map and --load-iterations on the CLI; dead --cooldown removed.
48ca5b3 to
6c69c9d
Compare
`yarn bench:docs` reads runs/<deviceId>/*.jsonl beside the docs page and writes measured.json, so the published numbers are derived from runs that are versioned with the docs and can be used as bench:compare baselines.
## Description - Benchmark rows move from `src/data/` (shared by every docs version) to beside the page, so a version cut archives each release's numbers. - The page now says the numbers are from 0.11.0-dev, not v0.10.0. - Rows come from `measured.json` (generated by `yarn bench:docs` from raw runs, #1363) merged with `imported.json` (current rows, no raw runs). Published numbers are unchanged. ### Introduces a breaking change? - [ ] Yes - [x] No ### Type of change - [ ] Bug fix (change which fixes an issue) - [ ] New feature (change which adds functionality) - [x] Documentation update (improves or adds clarity to existing documentation) - [ ] Other (chores, tests, code style improvements etc.) ### Tested on - [ ] iOS - [ ] Android ### Testing instructions ```bash cd docs && yarn build ``` `/docs/next/benchmarks` renders the charts. ### Screenshots ### Related issues #727, #1457 ### Checklist - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have updated the documentation accordingly - [x] My changes generate no new warnings ### Additional notes
barhanc
left a comment
There was a problem hiding this comment.
This is also missing documentation for the new core API.
| import { rnexecutorchJsi } from '../native/bridge'; | ||
|
|
||
| /** What one exported method accumulated since the last reset. */ | ||
| export interface MethodProfile { |
There was a problem hiding this comment.
Why interface and not type here? Also the jsdoc isn't great - not very descriptive and missing category tag.
| } | ||
| } // namespace | ||
|
|
||
| void record(const std::string &methodName, int64_t nanos) { |
There was a problem hiding this comment.
Imo key should also contain model path so e.g. two calls to forward on two different models aren't accumulated.
| entry.nanos += nanos; | ||
| } | ||
|
|
||
| void install_executionProfile(jsi::Runtime &rt, jsi::Object &module) { |
There was a problem hiding this comment.
Instead of setting these two on root module, set it on .profiler and maybe rename to install_profiler.
| const auto *name = "getExecutionProfile"; | ||
| auto fnBody = [](jsi::Runtime &rt, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, | ||
| size_t count) -> jsi::Value { | ||
| if (count != 0) { |
There was a problem hiding this comment.
This should imo allow passing an optional key so instead of profiler.getExecutionProfile()[key] one can write profiler.getExecutionProfile(key)
| // Core primitives — for library builders and power users | ||
| export * from './core/error'; | ||
| export * from './core/model'; | ||
| export * from './core/profiler'; |
There was a problem hiding this comment.
Maybe export * as profiler from ... and rename the methods to just get, reset so that idiomatic way to use is as follows
import { profiler } from 'react-native-executorch';
// ...
const stats = profiler.get(key);
profiler.reset();| * How much of a run was spent inside ExecuTorch. | ||
| * | ||
| * Every `execute` call adds its duration to a process-global tally, keyed by | ||
| * method name. Read it around a piece of work to learn what share of that work | ||
| * was the model, as opposed to the preprocessing and post-processing around it. | ||
| * | ||
| * This exists because the alternative does not work. A task pipeline owns its | ||
| * `Model` privately, so there is no handle to time; and re-running the model | ||
| * separately means guessing the shapes the pipeline fed it. For a model with a | ||
| * dynamic dimension the guess is wrong by construction: a text embedder | ||
| * declaring up to 510 tokens gets benchmarked at 510 while the pipeline ran 75, | ||
| * and the resulting "share" is meaningless. Accumulating in place removes the | ||
| * guess entirely, and it works for a pipeline that calls one method many times | ||
| * (OCR runs its recognizer once per detected box) where a single replay would not. |
There was a problem hiding this comment.
This doesn't really read like a public API doc, rather a comment explanation.
| auto fnBody = [](jsi::Runtime & /*rt*/, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, | ||
| size_t count) -> jsi::Value { |
There was a problem hiding this comment.
| auto fnBody = [](jsi::Runtime & /*rt*/, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, | |
| size_t count) -> jsi::Value { | |
| auto fnBody = [](jsi::Runtime & /*rt*/, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, size_t count) -> jsi::Value { |
| auto fnBody = [](jsi::Runtime &rt, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, | ||
| size_t count) -> jsi::Value { |
There was a problem hiding this comment.
| auto fnBody = [](jsi::Runtime &rt, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, | |
| size_t count) -> jsi::Value { | |
| auto fnBody = [](jsi::Runtime &rt, const jsi::Value & /*thisVal*/, const jsi::Value * /*args*/, size_t count) -> jsi::Value { |
| // Milliseconds as a double: the underlying clock is nanoseconds | ||
| // and a double holds that exactly well past any plausible run, | ||
| // so a sub-millisecond inference is not rounded to zero the way | ||
| // the integer-millisecond profiling log rounds it. |
There was a problem hiding this comment.
Does that actually require a comment?
| readonly totalMs: number; | ||
| } | ||
|
|
||
| /** Per-method totals, keyed by exported method name. */ |
Description
Adds
apps/benchmarks, a headless Expo app that measures load time, inference latency and peak memory for every published registry variant (208 Android / 255 iOS, derived frommodels.ts), plus a collector, a summarizer and a comparator.Notes:
getExecutionProfile()API, soExecute %reflects the shapes and call counts the pipeline actually used..ptepages.Protocol for other devices:
apps/benchmarks/BENCHMARK_SPEC.md.Introduces a breaking change?
Type of change
Tested on
Testing instructions
Screenshots
Related issues
Closes #1078
Checklist
Additional notes