Skip to content

feat(android): add the Qualcomm QNN backend - #1482

Draft
msluszniak wants to merge 11 commits into
mainfrom
@ms/qnn-backend
Draft

msluszniak wants to merge 11 commits into
mainfrom
@ms/qnn-backend

Conversation

@msluszniak

@msluszniak msluszniak commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Description

  • Opt-in qnn backend (Android arm64): libqnn_executorch_backend.so from v0.10.5-libs, Qualcomm runtime from Maven (com.qualcomm.qti:qnn-runtime:2.47.0).
  • A QNN .pte targets one Hexagon version, so QNN variants resolve per device via a new qnnHtpArch JSI value (ro.soc.model → v69–v81). QNN is a DEFAULT candidate only when it is set.
  • Declares libcdsprpc.so in the manifest and sets ADSP_LIBRARY_PATH; apps need useLegacyPackaging = true (documented).
  • EfficientNetV2-S gains QNN_A16W8 (S26 Ultra: 74.1% top-1 on ImageNetV2 vs 74.4% fp32, 2.4 ms vs 46.9 ms XNNPACK fp32).
  • EfficientNetV2-S QNN files re-exported with MinMaxObserver activations: 70.0% to 71.8% top-1 on 500 ImageNetV2 images (fp32 72.2%).
  • DeepLabV3-MobileNetV3 gains QNN_A16W8, also in the CV demo: 66.55% mIoU on VOC2012 vs 66.56% XNNPACK int8, 3.2 ms.
  • DeepLabV3-ResNet50 gains QNN_A16W8, also in the CV demo: 76.74% mIoU vs 76.71% XNNPACK int8, 20.1 ms.
  • DeepLabV3-ResNet101, FCN-ResNet50 and FCN-ResNet101 gain QNN_A16W8, also in the CV demo: 78.90% / 71.21% / 74.93% mIoU vs 78.82% / 71.36% / 74.99% XNNPACK int8.
  • LR-ASPP gains QNN_A16W8, also in the CV demo: 65.83% mIoU vs 65.06% XNNPACK int8, 2.1 ms.
  • DeepLabV3-MobileNetV3 QNN files re-exported with eps=1e-5: 66.55% to 69.02% mIoU (fp32 69.28%).

Stacked on #1479 (includes its commit until it merges).

Introduces a breaking change?

  • Yes
  • No

Type of change

  • Bug fix (change which fixes an issue)
  • New feature (change which adds functionality)
  • Documentation update (improves or adds clarity to existing documentation)
  • Other (chores, tests, code style improvements etc.)

Tested on

  • iOS
  • Android

Testing instructions

  • In apps/computer-vision/package.json add "backends": ["qnn"]; set expo.useLegacyPackaging=true in android/gradle.properties.
  • Run on a Snapdragon 8 Gen 1+ device, open Classification, pick QNN A16W8.

Screenshots

Related issues

Checklist

  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have updated the documentation accordingly
  • My changes generate no new warnings

Additional notes

Only the v81 file is device-validated; v69–v79 are untested.

…rsion

Adds the 44 Vulkan .pte files published under the HF v0.11.0 tag to the
registry: style transfer (4), YOLO26 detection (15), YOLO26-seg (15),
YOLO26-pose (3), RF-DETR detector/segmentation/keypoint, FastSAM s/x,
BlazeFace and SDXS. They become the Android default through the existing
`variants` resolution.

A 0.11 tag carries only the files cut under it, so the new URLs resolve
through NEXT_VERSION_TAG while each model's older exports stay on
VERSION_TAG.

FEATURE_MAP and the native-libraries doc gain `vulkan` for the features
that now reference it, so an app opting in by feature still downloads the
backend its models need.

nativeLibsVersion moves to 0.10.4, which is the first libs release whose
core carries the constant-segment fix these models need to load.
Opt-in `qnn` backend: links libqnn_executorch_backend.so from the libs
release and Qualcomm's runtime from Maven (com.qualcomm.qti:qnn-runtime
2.47.0). A QNN .pte targets one Hexagon version, so QNN variants resolve to
the device's file through the new `qnnHtpArch` JSI value (ro.soc.model ->
v69..v81), and QNN is only a DEFAULT candidate when that value is set.

EfficientNetV2-S gains QNN_A16W8. nativeLibsVersion moves to 0.10.5.
@msluszniak msluszniak self-assigned this Sep 22, 2026
Returns a class index per pixel rather than 21 logit planes, so the segmenter
takes its index-map path and skips the transpose and argmax over H*W*K floats.

66.55% mIoU on an S26 Ultra over 200 VOC2012 val images, against 66.56% for the
XNNPACK int8 file, at 3.20 ms/inference.
# Conflicts:
#	packages/react-native-executorch/package.json
#	packages/react-native-executorch/scripts/download-libs.js
#	packages/react-native-executorch/src/models.ts
76.74% mIoU on 200 VOC2012 val images on an S26 Ultra, against 76.71% for the
XNNPACK int8 file, at 20.1 ms/inference.
On an S26 Ultra over 200 VOC2012 val images, against the XNNPACK int8 files:
DeepLabV3-ResNet101 78.90% vs 78.82% mIoU, FCN-ResNet50 71.21% vs 71.36%,
FCN-ResNet101 74.93% vs 74.99%.
65.83% mIoU on an S26 Ultra over 200 VOC2012 val images, against 65.06% for the
XNNPACK int8 file, at 2.06 ms/inference.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant