Repository navigation
fix: backport the 0.10.3 library fixes (#1466, #1486, #1491, #1492, #1494) - #1497
Merged
Merged
Conversation
Fixes the parts of #1464 that hold up under checking. The opt-out gated downloads and linking; it did not control what the `core-*` artifacts contain, and on Android it did not control what ships. Artifacts republished as [`v0.10.1-libs`](https://github.com/software-mansion/react-native-executorch/releases/tag/v0.10.1-libs). **Repackaging only: every library byte is identical to v0.10.0-libs, nothing was recompiled**, so this cannot introduce a behaviour change or hit the Vulkan build's non-determinism. | artifact | v0.10.0-libs | v0.10.1-libs | |---|---|---| | core-android-arm64-v8a | 7.96 MB | 4.49 MB | | core-android-x86_64 | 8.50 MB | 4.50 MB | | core-ios | 20.82 MB | 5.87 MB | - **`package-release-artifacts.sh`** stages the runtime only into core. MPS goes with it (§9): `libbackend_mps_*.a` had no toggle in `ALL_BACKENDS` and the podspec never referenced it. - **`build.gradle.kts`** drives `packaging.jniLibs.excludes` from the resolved config (§4, §5), the Android counterpart of the podspec choosing its `-force_load` entries. - **`download-libs.js`** prunes backends absent from the config after extracting (§3, §6), so turning one off after an install takes effect instead of silently shipping the old binary. - **`COPYFILE_DISABLE=1`** when packing (§12). **§4 was inferred in the issue; it is now observed.** apps/legacy/bare-rn, release, arm64-v8a, `backends: ["vulkan"]`, old core staged so XNNPACK is present on disk: | build | merged_native_libs | libxnnpack | |---|---|---| | without the gradle fix | 269,127,104 B | present, 2,673,072 B | | with `jniLibs.excludes` | 266,454,032 B | absent | Delta is exactly the XNNPACK library. Read from `merged_native_libs/release/.../arm64-v8a`, the tree the APK packages. **iOS**: a Release link of the same app against the reduced `core-ios` succeeds. Everything the linker takes from `third-party/ios`: ``` libs/executorch/libthreadpool_ios.a XnnpackBackend.xcframework/ios-arm64/libXnnpackBackend.a (-force_load) CoreMLBackend.xcframework/ios-arm64/libCoreMLBackend.a (-force_load) MLXBackend.xcframework/ios-arm64/libMLXBackend.a (-force_load) + ExecutorchLib.framework ``` None of the 21 removed archives appears, so §5 is confirmed by the linker rather than by reading the podspec. **§6**: turning Vulkan on, then off, now removes the stale binary instead of leaving it to ship. **§12: there are no AppleDouble sidecars in v0.10.0-libs.** `tar tzf | grep -c '\._'` is 0 for all eight tarballs checked, against the 7 the issue reports. `COPYFILE_DISABLE=1` is added anyway, since it costs nothing and the failure mode is real on macOS. §7 root cause (the two Vulkan binaries differ; fixing §1 removes the *consequence*, since the standalone build is now the only one that can ship, but not the cause), §8 and §11 (inferred from `-force_load` and section sizes, no archive build run), §10 simulator split, §13 docs. - [ ] Yes - [x] No - [x] Bug fix (non-breaking change which fixes an issue) - [x] iOS - [x] Android `nativeLibsVersion` is bumped to 0.10.1, so a fresh install pulls the new artifacts. Declare `backends: ["vulkan"]` in an app's package.json, build a release APK, and check `unzip -l app-release.apk | grep 'lib/arm64-v8a/.*\.so'`: `libxnnpack_executorch_backend.so` should be absent. Flip a backend off after an install and confirm the postinstall log reports removing it. Refs #1464. --------- Co-authored-by: Mateusz Słuszniak <msluszniak1@gmail.com>
## Description A `.pte` declares its own stop token ids, and they routinely outnumber `eos_token`. Qwen ends a turn with `<|im_end|>` but also stops on `<|endoftext|>`, which the tokenizer config only names as `pad_token`, so that one reached the user verbatim. 0.9 stripped `eos_token`, `eot_token` and `pad_token`; the chat session only dropped the first. - `parseTokenizerConfig` resolves `eot_token` and `pad_token` too and returns them as `stopTokens` (eos first, deduped). `eosToken` stays, so the change is additive. - The generation callback tests each chunk against the whole stop list rather than against `eos_token` alone. ExecuTorch decodes one token per callback, so a membership test covers it. Affects qwen3, qwen2.5 and hammer2.1, which all have `pad_token` different from `eos_token`. ### Introduces a breaking change? - [ ] Yes - [x] No ### Type of change - [x] Bug fix (change which fixes an issue) - [ ] New feature (change which adds functionality) - [ ] Documentation update (improves or adds clarity to existing documentation) - [ ] Other (chores, tests, code style improvements etc.) ### Tested on - [ ] iOS - [ ] Android ### Testing instructions - `yarn jest __tests__/tasks/llmChatSession.test.ts` in `packages/react-native-executorch`. - Two new cases: a terminal token named only as `pad_token`, and one named only as `eot_token`. Both fail on the current source and pass with this change. ### Screenshots ### Related issues ### Checklist - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have updated the documentation accordingly - [x] My changes generate no new warnings ### Additional notes Not run on a device: the change is TS only and the behaviour is covered by the unit tests above. Found while looking into Qwen3 output in software-mansion-labs/react-native-rag#23, where the adapter needed a `stopRegex: /<\|endoftext\|>/` to work around it. The other half of that report was the Qwen3 `chat_template`, which was byte identical to Qwen2.5's on the HF repo and is fixed there (`v0.10.0` now points at `5004436`), so it needs no change here. --------- Co-authored-by: Bartosz Hanc <bartosz.hanc02@gmail.com>
…ults (#1492) ## Description Neither factory was called with a load mode, so both inherited upstream's, and the two defaults are wrong in different ways. `create_multimodal_runner` defaults to `LoadMode::File`, which reads the whole `.pte` into a heap buffer. With `gemma4_e2b_mlx_int4.pte` (2.9 GB) that walks `phys_footprint` up in ~93 MB steps to 3296 MB and jetsam kills the app during load. `create_text_llm_runner` defaults to `MmapUseMlockIgnoreErrors`, which leaves footprint alone but wires the model resident. That is the number `DeviceInfo.getUsedMemory()` and every other memory API reachable from JS reports, so the new API looked like it used 6x the memory of the legacy one on the same model. Pin `Mmap` for both, which is what the legacy binding has always passed. ### Introduces a breaking change? - [ ] Yes - [x] No ### Type of change - [x] Bug fix (change which fixes an issue) - [ ] New feature (change which adds functionality) - [ ] Documentation update (improves or adds clarity to existing documentation) - [ ] Other (chores, tests, code style improvements etc.) ### Tested on - [x] iOS - [ ] Android ### Testing instructions iPhone 16 Release, gemma-4-e2b-mlx, after `load()`: | load mode | footprint | resident | | --- | --- | --- | | `File` | dies climbing past 3296 MB | n/a | | `MmapUseMlockIgnoreErrors` (old text default) | 930-954 MB | 3135-3282 MB | | `Mmap` (this PR, and legacy) | 950-952 MB | 355-475 MB | ### Related issues #1489 ### Checklist - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have updated the documentation accordingly - [x] My changes generate no new warnings ### Additional notes Neither mode changes what jetsam sees on the text path, so this is not the fix for #1489's prefill growth. It is what makes the new API's reported memory match the legacy API's.
## Description Upstream sizes prefill chunks from `get_max_seq_len`, which is the KV context budget rather than the widest tensor the graph accepts. `TextPrefiller` then never chunks, and any prompt past the real bound fails in `internal_resize_contiguous` with `Error::NotSupported`. Both gemma4 builds with a sliding window disagree: | model | `get_max_seq_len` | real `forward` bound | | --- | --- | --- | | `gemma4_e2b_mlx_int4` | 2048 | 511 | | `gemma_4_e2b_xnnpack_8da4w` | 4048 | 1024 | We now read the real bound from `method_meta(...).input_tensor_meta(0)` and lower the prefiller's chunk size to it, which is what the legacy runner has always done. `MultimodalPrefiller` has no chunking at all, so its text inputs are split into `TOKENS` pieces instead. ### Introduces a breaking change? - [ ] Yes - [x] No ### Type of change - [x] Bug fix (change which fixes an issue) - [ ] New feature (change which adds functionality) - [ ] Documentation update (improves or adds clarity to existing documentation) - [ ] Other (chores, tests, code style improvements etc.) ### Tested on - [x] iOS - [ ] Android ### Testing instructions `apps/nlp`, gemma-4-e2b-mlx, iPhone 16 Release. Send a prompt over 511 tokens (~6000 chars). A/B on the same binary: | prompt | before | after | | --- | --- | --- | | 407 tokens | ok | ok | | 1012 tokens | `Error::NotSupported`, KV position does not move | ok, 1495 MB peak, coherent output | ### Related issues #1489 ### Checklist - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have updated the documentation accordingly - [x] My changes generate no new warnings ### Additional notes Not gated on a backend. The clamp is a no-op when the two numbers agree, which is the case for the Vulkan gemma4 build (128/128) and for every other LLM checked, where the bound is just `max_seq_len - 1`. The legacy runner reads the same metadata but only when `uses_backend("MLXBackend")`, so it is still exposed on gemma4 XNNPACK.
## Description The postinstall downloader shelled out through shell strings twice, and both broke on Windows. - `sha256()` used `sha256sum`, which prefixes its whole output line with a backslash when the filename holds one, so every hash of a `C:\...` cache path came back as `\<hex>` and no artifact validated. Now hashed in-process with `crypto`, chunked. - `extract()` passed the tarball as `-f <path>`, which GNU tar reads as `host:file` and tries to resolve a host named `C`. The archive now arrives on stdin and `tar` runs with `cwd` set to the target dir, since GNU tar also unquotes `-C` paths (`\r`, `\t` in `node_modules\react-native-executorch\third-party`). A missing or failing `tar` now says which archive and why. - Also documents `tar` as a requirement and that Windows/Linux hosts build for Android only. ### Introduces a breaking change? - [ ] Yes - [x] No ### Type of change - [x] Bug fix (change which fixes an issue) - [ ] New feature (change which adds functionality) - [ ] Documentation update (improves or adds clarity to existing documentation) - [ ] Other (chores, tests, code style improvements etc.) ### Tested on - [x] iOS - [x] Android ### Testing instructions - Full postinstall run against `v0.10.4-libs`: all 13 artifacts downloaded, verified and extracted; second run reports 13 cache hits. - New `__tests__/api/downloadLibs.test.ts` covers both regressions with backslash-bearing paths, plus a >1 MiB file and the failure message. ### Screenshots ### Related issues Fixes #1493 ### Checklist - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have updated the documentation accordingly - [x] My changes generate no new warnings ### Additional notes Verified on Windows 11 with GNU tar (Git for Windows) and bsdtar. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
v0.10.4-libs carries the XNNPACK weights-cache fixes: the PReLU use-after-free and the retained unpacked weights that got Kokoro XNNPACK killed on iPhone.
Merged
8 of 12 tasks
barhanc
approved these changes
Sep 25, 2026
msluszniak
added a commit
that referenced
this pull request
Sep 25, 2026
## Description Bumps the core package to v0.10.3. Stacked on #1497, so review them together; after #1497 merges this gets rebased onto `release/0.10`. `LIB_VERSION` in `src/fetcher/telemetry.ts` moves with `package.json`. The adapter packages and webrtc are untouched, so only the core publish workflow applies. ### Introduces a breaking change? - [ ] Yes - [x] No ### Type of change - [ ] Bug fix (change which fixes an issue) - [ ] New feature (change which adds functionality) - [ ] Documentation update (improves or adds clarity to existing documentation) - [x] Other (chores, tests, code style improvements etc.) ### Tested on - [x] iOS - [x] Android ### Testing instructions `yarn jest __tests__` in `packages/react-native-executorch`. ### Screenshots ### Related issues #1497 ### Checklist - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have updated the documentation accordingly - [x] My changes generate no new warnings ### Additional notes
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Cherry-picks five fixes from
mainontorelease/0.10for the v0.10.3 patch, and pinsnativeLibsVersionto0.10.4.fix(libs): stop shipping backends the app opted out of(fix(libs): stop shipping backends the app opted out of #1466)fix(llm): keep every terminal token out of the chat response(fix(llm): keep every terminal token out of the chat response #1486)fix(llm): pin the model load mode instead of inheriting upstream defaults(fix(llm): pin the model load mode instead of inheriting upstream defaults #1492)fix(llm): chunk prefill at the bound the graph declares(fix(llm): chunk prefill at the bound the graph declares #1491)fix(install): make the native lib download work on Windows(fix(install): make the native lib download work on Windows #1494)v0.10.4-libscarries the XNNPACK weights-cache fixes (PReLU use-after-free, and the retained weights that got Kokoro XNNPACK killed on iPhone). Docs hunks from #1466 and #1494 are left out, per the patch-release rule. Every touched file matchesmainexceptpackage.jsonanddownload-libs.js, which lacks #1479's Vulkan task map.Introduces a breaking change?
Type of change
Tested on
Testing instructions
yarn jest __tests__,yarn typecheck,yarn lint,scripts/run-native-tests.shapps/speechRelease on iPhone 16 and Galaxy S26 Ultra: Kokoro EN_US XNNPACK synthesizes, and a 12-layer PReLU.ptematches eager output.Screenshots
Related issues
Checklist
Additional notes
The version bump follows in a separate
Release v0.10.3PR once this merges.