Skip to content

fix: backport the 0.10.3 library fixes (#1466, #1486, #1491, #1492, #1494) - #1497

Merged
msluszniak merged 6 commits into
release/0.10from
@ms/backport-0.10.3
Sep 25, 2026
Merged

msluszniak merged 6 commits into
release/0.10from
@ms/backport-0.10.3

Conversation

@msluszniak

Copy link
Copy Markdown
Member

Description

Cherry-picks five fixes from main onto release/0.10 for the v0.10.3 patch, and pins nativeLibsVersion to 0.10.4.

v0.10.4-libs carries the XNNPACK weights-cache fixes (PReLU use-after-free, and the retained weights that got Kokoro XNNPACK killed on iPhone). Docs hunks from #1466 and #1494 are left out, per the patch-release rule. Every touched file matches main except package.json and download-libs.js, which lacks #1479's Vulkan task map.

Introduces a breaking change?

  • Yes
  • No

Type of change

  • Bug fix (change which fixes an issue)
  • New feature (change which adds functionality)
  • Documentation update (improves or adds clarity to existing documentation)
  • Other (chores, tests, code style improvements etc.)

Tested on

  • iOS
  • Android

Testing instructions

  • yarn jest __tests__, yarn typecheck, yarn lint, scripts/run-native-tests.sh
  • apps/speech Release on iPhone 16 and Galaxy S26 Ultra: Kokoro EN_US XNNPACK synthesizes, and a 12-layer PReLU .pte matches eager output.

Screenshots

Related issues

Checklist

  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have updated the documentation accordingly
  • My changes generate no new warnings

Additional notes

The version bump follows in a separate Release v0.10.3 PR once this merges.

msluszniak and others added 6 commits September 25, 2026 18:53
Fixes the parts of #1464 that hold up under checking. The opt-out gated
downloads and linking; it did not control what the `core-*` artifacts
contain, and on Android it did not control what ships.

Artifacts republished as
[`v0.10.1-libs`](https://github.com/software-mansion/react-native-executorch/releases/tag/v0.10.1-libs).
**Repackaging only: every library byte is identical to v0.10.0-libs,
nothing was recompiled**, so this cannot introduce a behaviour change or
hit the Vulkan build's non-determinism.

| artifact | v0.10.0-libs | v0.10.1-libs |
|---|---|---|
| core-android-arm64-v8a | 7.96 MB | 4.49 MB |
| core-android-x86_64 | 8.50 MB | 4.50 MB |
| core-ios | 20.82 MB | 5.87 MB |

- **`package-release-artifacts.sh`** stages the runtime only into core.
MPS goes with it (§9): `libbackend_mps_*.a` had no toggle in
`ALL_BACKENDS` and the podspec never referenced it.
- **`build.gradle.kts`** drives `packaging.jniLibs.excludes` from the
resolved config (§4, §5), the Android counterpart of the podspec
choosing its `-force_load` entries.
- **`download-libs.js`** prunes backends absent from the config after
extracting (§3, §6), so turning one off after an install takes effect
instead of silently shipping the old binary.
- **`COPYFILE_DISABLE=1`** when packing (§12).

**§4 was inferred in the issue; it is now observed.**
apps/legacy/bare-rn, release, arm64-v8a, `backends: ["vulkan"]`, old
core staged so XNNPACK is present on disk:

| build | merged_native_libs | libxnnpack |
|---|---|---|
| without the gradle fix | 269,127,104 B | present, 2,673,072 B |
| with `jniLibs.excludes` | 266,454,032 B | absent |

Delta is exactly the XNNPACK library. Read from
`merged_native_libs/release/.../arm64-v8a`, the tree the APK packages.

**iOS**: a Release link of the same app against the reduced `core-ios`
succeeds. Everything the linker takes from `third-party/ios`:

```
libs/executorch/libthreadpool_ios.a
XnnpackBackend.xcframework/ios-arm64/libXnnpackBackend.a   (-force_load)
CoreMLBackend.xcframework/ios-arm64/libCoreMLBackend.a     (-force_load)
MLXBackend.xcframework/ios-arm64/libMLXBackend.a           (-force_load)
+ ExecutorchLib.framework
```

None of the 21 removed archives appears, so §5 is confirmed by the
linker rather than by reading the podspec.

**§6**: turning Vulkan on, then off, now removes the stale binary
instead of leaving it to ship.

**§12: there are no AppleDouble sidecars in v0.10.0-libs.** `tar tzf |
grep -c '\._'` is 0 for all eight tarballs checked, against the 7 the
issue reports. `COPYFILE_DISABLE=1` is added anyway, since it costs
nothing and the failure mode is real on macOS.

§7 root cause (the two Vulkan binaries differ; fixing §1 removes the
*consequence*, since the standalone build is now the only one that can
ship, but not the cause), §8 and §11 (inferred from `-force_load` and
section sizes, no archive build run), §10 simulator split, §13 docs.

- [ ] Yes
- [x] No

- [x] Bug fix (non-breaking change which fixes an issue)

- [x] iOS
- [x] Android

`nativeLibsVersion` is bumped to 0.10.1, so a fresh install pulls the
new artifacts. Declare `backends: ["vulkan"]` in an app's package.json,
build a release APK, and check `unzip -l app-release.apk | grep
'lib/arm64-v8a/.*\.so'`: `libxnnpack_executorch_backend.so` should be
absent. Flip a backend off after an install and confirm the postinstall
log reports removing it.

Refs #1464.

---------

Co-authored-by: Mateusz Słuszniak <msluszniak1@gmail.com>
## Description

A `.pte` declares its own stop token ids, and they routinely outnumber
`eos_token`. Qwen ends a turn with `<|im_end|>` but also stops on
`<|endoftext|>`, which the tokenizer config only names as `pad_token`,
so that one reached the user verbatim. 0.9 stripped `eos_token`,
`eot_token` and `pad_token`; the chat session only dropped the first.

- `parseTokenizerConfig` resolves `eot_token` and `pad_token` too and
returns them as `stopTokens` (eos first, deduped). `eosToken` stays, so
the change is additive.
- The generation callback tests each chunk against the whole stop list
rather than against `eos_token` alone. ExecuTorch decodes one token per
callback, so a membership test covers it.

Affects qwen3, qwen2.5 and hammer2.1, which all have `pad_token`
different from `eos_token`.

### Introduces a breaking change?

- [ ] Yes
- [x] No

### Type of change

- [x] Bug fix (change which fixes an issue)
- [ ] New feature (change which adds functionality)
- [ ] Documentation update (improves or adds clarity to existing
documentation)
- [ ] Other (chores, tests, code style improvements etc.)

### Tested on

- [ ] iOS
- [ ] Android

### Testing instructions

- `yarn jest __tests__/tasks/llmChatSession.test.ts` in
`packages/react-native-executorch`.
- Two new cases: a terminal token named only as `pad_token`, and one
named only as `eot_token`. Both fail on the current source and pass with
this change.

### Screenshots

### Related issues

### Checklist

- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have updated the documentation accordingly
- [x] My changes generate no new warnings

### Additional notes

Not run on a device: the change is TS only and the behaviour is covered
by the unit tests above.

Found while looking into Qwen3 output in
software-mansion-labs/react-native-rag#23, where the adapter needed a
`stopRegex: /<\|endoftext\|>/` to work around it. The other half of that
report was the Qwen3 `chat_template`, which was byte identical to
Qwen2.5's on the HF repo and is fixed there (`v0.10.0` now points at
`5004436`), so it needs no change here.

---------

Co-authored-by: Bartosz Hanc <bartosz.hanc02@gmail.com>
…ults (#1492)

## Description

Neither factory was called with a load mode, so both inherited
upstream's, and the two defaults are wrong in different ways.

`create_multimodal_runner` defaults to `LoadMode::File`, which reads the
whole `.pte` into a heap buffer. With `gemma4_e2b_mlx_int4.pte` (2.9 GB)
that walks `phys_footprint` up in ~93 MB steps to 3296 MB and jetsam
kills the app during load.

`create_text_llm_runner` defaults to `MmapUseMlockIgnoreErrors`, which
leaves footprint alone but wires the model resident. That is the number
`DeviceInfo.getUsedMemory()` and every other memory API reachable from
JS reports, so the new API looked like it used 6x the memory of the
legacy one on the same model.

Pin `Mmap` for both, which is what the legacy binding has always passed.

### Introduces a breaking change?

- [ ] Yes
- [x] No

### Type of change

- [x] Bug fix (change which fixes an issue)
- [ ] New feature (change which adds functionality)
- [ ] Documentation update (improves or adds clarity to existing
documentation)
- [ ] Other (chores, tests, code style improvements etc.)

### Tested on

- [x] iOS
- [ ] Android

### Testing instructions

iPhone 16 Release, gemma-4-e2b-mlx, after `load()`:

| load mode | footprint | resident |
| --- | --- | --- |
| `File` | dies climbing past 3296 MB | n/a |
| `MmapUseMlockIgnoreErrors` (old text default) | 930-954 MB | 3135-3282
MB |
| `Mmap` (this PR, and legacy) | 950-952 MB | 355-475 MB |

### Related issues

#1489

### Checklist

- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have updated the documentation accordingly
- [x] My changes generate no new warnings

### Additional notes

Neither mode changes what jetsam sees on the text path, so this is not
the fix for #1489's prefill growth. It is what makes the new API's
reported memory match the legacy API's.
## Description

Upstream sizes prefill chunks from `get_max_seq_len`, which is the KV
context budget rather than the widest tensor the graph accepts.
`TextPrefiller` then never chunks, and any prompt past the real bound
fails in `internal_resize_contiguous` with `Error::NotSupported`. Both
gemma4 builds with a sliding window disagree:

| model | `get_max_seq_len` | real `forward` bound |
| --- | --- | --- |
| `gemma4_e2b_mlx_int4` | 2048 | 511 |
| `gemma_4_e2b_xnnpack_8da4w` | 4048 | 1024 |

We now read the real bound from `method_meta(...).input_tensor_meta(0)`
and lower the prefiller's chunk size to it, which is what the legacy
runner has always done. `MultimodalPrefiller` has no chunking at all, so
its text inputs are split into `TOKENS` pieces instead.

### Introduces a breaking change?

- [ ] Yes
- [x] No

### Type of change

- [x] Bug fix (change which fixes an issue)
- [ ] New feature (change which adds functionality)
- [ ] Documentation update (improves or adds clarity to existing
documentation)
- [ ] Other (chores, tests, code style improvements etc.)

### Tested on

- [x] iOS
- [ ] Android

### Testing instructions

`apps/nlp`, gemma-4-e2b-mlx, iPhone 16 Release. Send a prompt over 511
tokens (~6000 chars). A/B on the same binary:

| prompt | before | after |
| --- | --- | --- |
| 407 tokens | ok | ok |
| 1012 tokens | `Error::NotSupported`, KV position does not move | ok,
1495 MB peak, coherent output |

### Related issues

#1489

### Checklist

- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have updated the documentation accordingly
- [x] My changes generate no new warnings

### Additional notes

Not gated on a backend. The clamp is a no-op when the two numbers agree,
which is the case for the Vulkan gemma4 build (128/128) and for every
other LLM checked, where the bound is just `max_seq_len - 1`.

The legacy runner reads the same metadata but only when
`uses_backend("MLXBackend")`, so it is still exposed on gemma4 XNNPACK.
## Description

The postinstall downloader shelled out through shell strings twice, and
both broke on Windows.

- `sha256()` used `sha256sum`, which prefixes its whole output line with
a backslash when the filename holds one, so every hash of a `C:\...`
cache path came back as `\<hex>` and no artifact validated. Now hashed
in-process with `crypto`, chunked.
- `extract()` passed the tarball as `-f <path>`, which GNU tar reads as
`host:file` and tries to resolve a host named `C`. The archive now
arrives on stdin and `tar` runs with `cwd` set to the target dir, since
GNU tar also unquotes `-C` paths (`\r`, `\t` in
`node_modules\react-native-executorch\third-party`). A missing or
failing `tar` now says which archive and why.
- Also documents `tar` as a requirement and that Windows/Linux hosts
build for Android only.

### Introduces a breaking change?

- [ ] Yes
- [x] No

### Type of change

- [x] Bug fix (change which fixes an issue)
- [ ] New feature (change which adds functionality)
- [ ] Documentation update (improves or adds clarity to existing
documentation)
- [ ] Other (chores, tests, code style improvements etc.)

### Tested on

- [x] iOS
- [x] Android

### Testing instructions

- Full postinstall run against `v0.10.4-libs`: all 13 artifacts
downloaded, verified and extracted; second run reports 13 cache hits.
- New `__tests__/api/downloadLibs.test.ts` covers both regressions with
backslash-bearing paths, plus a >1 MiB file and the failure message.

### Screenshots

### Related issues

Fixes #1493

### Checklist

- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have updated the documentation accordingly
- [x] My changes generate no new warnings

### Additional notes

Verified on Windows 11 with GNU tar (Git for Windows) and bsdtar.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
v0.10.4-libs carries the XNNPACK weights-cache fixes: the PReLU use-after-free
and the retained unpacked weights that got Kokoro XNNPACK killed on iPhone.
@msluszniak
msluszniak requested a review from barhanc September 25, 2026 17:20
@msluszniak msluszniak self-assigned this Sep 25, 2026
@msluszniak msluszniak added chore PRs that are chores bug fix PRs that are fixing bugs labels Sep 25, 2026
@msluszniak msluszniak mentioned this pull request Sep 25, 2026
8 of 12 tasks
@msluszniak
msluszniak merged commit a04c463 into release/0.10 Sep 25, 2026
@msluszniak
msluszniak deleted the @ms/backport-0.10.3 branch September 25, 2026 17:36
msluszniak added a commit that referenced this pull request Sep 25, 2026
## Description

Bumps the core package to v0.10.3. Stacked on #1497, so review them
together; after #1497 merges this gets rebased onto `release/0.10`.

`LIB_VERSION` in `src/fetcher/telemetry.ts` moves with `package.json`.
The adapter packages and webrtc are untouched, so only the core publish
workflow applies.

### Introduces a breaking change?

- [ ] Yes
- [x] No

### Type of change

- [ ] Bug fix (change which fixes an issue)
- [ ] New feature (change which adds functionality)
- [ ] Documentation update (improves or adds clarity to existing
documentation)
- [x] Other (chores, tests, code style improvements etc.)

### Tested on

- [x] iOS
- [x] Android

### Testing instructions

`yarn jest __tests__` in `packages/react-native-executorch`.

### Screenshots

### Related issues

#1497

### Checklist

- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have updated the documentation accordingly
- [x] My changes generate no new warnings

### Additional notes
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug fix PRs that are fixing bugs chore PRs that are chores

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants