feat(android): add the Qualcomm QNN backend - #1482
Draft
msluszniak wants to merge 11 commits into
Draft
msluszniak wants to merge 11 commits into
msluszniak wants to merge 11 commits into
Conversation
…rsion Adds the 44 Vulkan .pte files published under the HF v0.11.0 tag to the registry: style transfer (4), YOLO26 detection (15), YOLO26-seg (15), YOLO26-pose (3), RF-DETR detector/segmentation/keypoint, FastSAM s/x, BlazeFace and SDXS. They become the Android default through the existing `variants` resolution. A 0.11 tag carries only the files cut under it, so the new URLs resolve through NEXT_VERSION_TAG while each model's older exports stay on VERSION_TAG. FEATURE_MAP and the native-libraries doc gain `vulkan` for the features that now reference it, so an app opting in by feature still downloads the backend its models need. nativeLibsVersion moves to 0.10.4, which is the first libs release whose core carries the constant-segment fix these models need to load.
Opt-in `qnn` backend: links libqnn_executorch_backend.so from the libs release and Qualcomm's runtime from Maven (com.qualcomm.qti:qnn-runtime 2.47.0). A QNN .pte targets one Hexagon version, so QNN variants resolve to the device's file through the new `qnnHtpArch` JSI value (ro.soc.model -> v69..v81), and QNN is only a DEFAULT candidate when that value is set. EfficientNetV2-S gains QNN_A16W8. nativeLibsVersion moves to 0.10.5.
Returns a class index per pixel rather than 21 logit planes, so the segmenter takes its index-map path and skips the transpose and argmax over H*W*K floats. 66.55% mIoU on an S26 Ultra over 200 VOC2012 val images, against 66.56% for the XNNPACK int8 file, at 3.20 ms/inference.
# Conflicts: # packages/react-native-executorch/package.json # packages/react-native-executorch/scripts/download-libs.js # packages/react-native-executorch/src/models.ts
76.74% mIoU on 200 VOC2012 val images on an S26 Ultra, against 76.71% for the XNNPACK int8 file, at 20.1 ms/inference.
On an S26 Ultra over 200 VOC2012 val images, against the XNNPACK int8 files: DeepLabV3-ResNet101 78.90% vs 78.82% mIoU, FCN-ResNet50 71.21% vs 71.36%, FCN-ResNet101 74.93% vs 74.99%.
65.83% mIoU on an S26 Ultra over 200 VOC2012 val images, against 65.06% for the XNNPACK int8 file, at 2.06 ms/inference.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
qnnbackend (Android arm64):libqnn_executorch_backend.sofromv0.10.5-libs, Qualcomm runtime from Maven (com.qualcomm.qti:qnn-runtime:2.47.0)..ptetargets one Hexagon version, so QNN variants resolve per device via a newqnnHtpArchJSI value (ro.soc.model→ v69–v81). QNN is aDEFAULTcandidate only when it is set.libcdsprpc.soin the manifest and setsADSP_LIBRARY_PATH; apps needuseLegacyPackaging = true(documented).QNN_A16W8(S26 Ultra: 74.1% top-1 on ImageNetV2 vs 74.4% fp32, 2.4 ms vs 46.9 ms XNNPACK fp32).MinMaxObserveractivations: 70.0% to 71.8% top-1 on 500 ImageNetV2 images (fp32 72.2%).QNN_A16W8, also in the CV demo: 66.55% mIoU on VOC2012 vs 66.56% XNNPACK int8, 3.2 ms.QNN_A16W8, also in the CV demo: 76.74% mIoU vs 76.71% XNNPACK int8, 20.1 ms.QNN_A16W8, also in the CV demo: 78.90% / 71.21% / 74.93% mIoU vs 78.82% / 71.36% / 74.99% XNNPACK int8.QNN_A16W8, also in the CV demo: 65.83% mIoU vs 65.06% XNNPACK int8, 2.1 ms.eps=1e-5: 66.55% to 69.02% mIoU (fp32 69.28%).Stacked on #1479 (includes its commit until it merges).
Introduces a breaking change?
Type of change
Tested on
Testing instructions
apps/computer-vision/package.jsonadd"backends": ["qnn"]; setexpo.useLegacyPackaging=trueinandroid/gradle.properties.QNN A16W8.Screenshots
Related issues
Checklist
Additional notes
Only the v81 file is device-validated; v69–v79 are untested.