Vulkan LLMs already work, but the win is TTFT and memory, not decode. Recording what would change that.
Measured, Qwen3-0.6B vs published XNNPACK:
|
Adreno 840 |
Mali-G76 |
| prefill |
2.7x faster |
4.5x slower |
| decode |
0.86x |
1.7x faster, 3.4x at 773 tokens |
| RSS |
303 vs 1088 MiB |
1038 vs 1028 MiB |
Decode is batch-1 GEMV, so it is bandwidth bound and the GPU has no FLOPs to spend. That explains both columns, including why Mali inverts: prefill per token there is about equal to decode per token, so batching buys nothing.
1. 4-bit weights on GPU
The biggest lever. Decode is bandwidth bound, so weight precision is the bottleneck. int8 is fixed (pytorch/executorch#22429, #22430) and measured faster than fp16 on the mpnet embedders. linear_qcs4w is broken the same way and untouched, so a Vulkan LLM decodes fp16 weights while XNNPACK decodes 4-bit. That is roughly a 4x bandwidth advantage to XNNPACK in exactly the phase Vulkan loses.
### 2. Adreno nondeterminism (pytorch/executorch#22327, #21938)
Correctness, not performance. One nn.Linear(384,384) at M=1500 is wrong on 23/100 executions, error 3x the signal, nothing reported. It is why LFM2.5-350M emitted garbage from a byte-identical cached file. Not root caused, no PR.
3. One artifact does not fit both vendors
force_fp16 is +14% decode on Adreno and loses on both axes on Mali. Vulkan int8 tiled kernels are Adreno tuned and pathological on Mali (5.4x on lfm2.5-embedding, 5-8x on mpnet). Needs per-vendor variant selection or half our users get a regression.
4. Dispatch overhead and occupancy
reduce_gwg hardcodes 16 threads, softmax_buffer hardcodes 4. LFM2.5 text TTFT was 80 vs 56 ms on short prompts, mostly fixed per-call cost.
Known good recipe
Keep for any new Vulkan LLM export: plain VulkanPartitioner so causal_sdpa fuses (57 delegates to 1, decode 41.5 to 86.2 tok/s), and enable_dynamic_shape=True, without which ET prefills one token at a time.
Vulkan LLMs already work, but the win is TTFT and memory, not decode. Recording what would change that.
Measured, Qwen3-0.6B vs published XNNPACK:
Decode is batch-1 GEMV, so it is bandwidth bound and the GPU has no FLOPs to spend. That explains both columns, including why Mali inverts: prefill per token there is about equal to decode per token, so batching buys nothing.
1. 4-bit weights on GPU
The biggest lever. Decode is bandwidth bound, so weight precision is the bottleneck. int8 is fixed (pytorch/executorch#22429, #22430) and measured faster than fp16 on the mpnet embedders.
linear_qcs4wis broken the same way and untouched, so a Vulkan LLM decodes fp16 weights while XNNPACK decodes 4-bit. That is roughly a 4x bandwidth advantage to XNNPACK in exactly the phase Vulkan loses.### 2. Adreno nondeterminism (pytorch/executorch#22327, #21938)Correctness, not performance. Onenn.Linear(384,384)at M=1500 is wrong on 23/100 executions, error 3x the signal, nothing reported. It is why LFM2.5-350M emitted garbage from a byte-identical cached file. Not root caused, no PR.3. One artifact does not fit both vendors
force_fp16is +14% decode on Adreno and loses on both axes on Mali. Vulkan int8 tiled kernels are Adreno tuned and pathological on Mali (5.4x on lfm2.5-embedding, 5-8x on mpnet). Needs per-vendor variant selection or half our users get a regression.4. Dispatch overhead and occupancy
reduce_gwghardcodes 16 threads,softmax_bufferhardcodes 4. LFM2.5 text TTFT was 80 vs 56 ms on short prompts, mostly fixed per-call cost.Known good recipe
Keep for any new Vulkan LLM export: plain
VulkanPartitionersocausal_sdpafuses (57 delegates to 1, decode 41.5 to 86.2 tok/s), andenable_dynamic_shape=True, without which ET prefills one token at a time.