Repository navigation
[Feature Request] Allow VAD to run on a different device than the ASR model (Apple Silicon MPS regression: VAD 5x slower than CPU) #3701
Description
Activity
Thanks for the detailed report and the offer to contribute a PR. I recommend honoring the existing nested device options rather than adding a second family of
vad_device/punc_device/spk_devicearguments. The proposed interface is:# Proposed behavior, not supported correctly by current main yet. model = AutoModel( model=asr, device="mps", vad_model=vad, vad_kwargs={"device": "cpu"}, punc_model=punc, punc_kwargs={"device": "cpu"}, )
I checked current main
486b4b7ceb27b72db24a84c6e1be7cf6c5be6609. A CPU-only configuration probe using the actual AutoModel orchestration and stand-in models confirmed three relevant paths:- Construction overwrites an explicit
device="cpu"in all three submodel dictionaries. It also modifies the caller's dictionaries in place. - Even after establishing a CPU VAD baseline, a shared
generate(device="cuda:0")config changes the device passed to VAD back tocuda:0; the next call without that override resets to CPU. This is thedeep_updatepath in bothinference_with_vadandinference, so changing only the constructor assignments would not cover it. - The speaker-embedding call at auto_model.py:982 receives the ASR
kwargs, notself.spk_kwargs. In the probe it received the ASR device and lost a speaker-only sentinel setting, despite a CPU speaker baseline.
For a PR, please make the contract consistent across VAD, punctuation and speaker models:
- An explicitly supplied submodel device takes precedence; when omitted, retain the existing inheritance from the resolved ASR device. Copy caller-owned config dictionaries before filling defaults.
- Treat the selected submodel device as model placement, not an ordinary shared per-call option. Preserve it through runtime config merging and repeated-call resets, keeping weights and input features on the same device. Do not imply that changing a runtime
devicestring moves already-loaded weights; runtime migration would need a separate explicit contract. - Pass the speaker model its own resolved config. Cover punctuation both with and without VAD, since both routes merge shared runtime options.
- Add regressions for explicit overrides, omitted-device inheritance, unchanged caller dictionaries, shared runtime options, two successive calls, and the actual per-model inference boundaries. Please also exercise unavailable-device fallback without claiming hardware execution from mocked availability.
A PR covering this would be welcome. Please include an actual mixed-device smoke test where hardware is available, checking model/feature device agreement and successful output. Your original Apple M2 audio is particularly valuable for validating the reported performance after the fix. My probe used CPU tensors only (
cuda:0was a configuration label), so I have not independently reproduced the MPS timing figures or a hardware device-mismatch crash. NoComputeScoresmonkeypatch should be required by the final public API.- Construction overwrites an explicit
Update: the nested-device fix has now been merged in PR #3706, at commit 4d8e087. The earlier comment's “not supported correctly by current main yet” note describes the old base and no longer applies to this merged source.
The public configuration is documented in the merged English and Chinese API guides:
- An explicit nested device in
vad_kwargs,punc_kwargs, orspk_kwargsis preserved; omitted values inherit the resolved ASR device. - Shared per-call options do not migrate already-loaded models.
- Caller-owned submodel dictionaries are copied, and the speaker model receives its own configuration.
I checked that all six merged PR files match the final candidate. Compared with the previously reviewed
eee63614, the only follow-up change is an optional-clustering test-fixture adjustment; the production AutoModel source is unchanged.This is a source merge, not a new PyPI release announcement. The original M2/121-minute performance case has not been independently retested, and the author's M4 measurements remain workload-specific. Keeping this issue open for original-environment confirmation; no internal VAD replacement or
ComputeScorespatch should be necessary for the documented placement configuration.- An explicit nested device in
Release availability update: the mixed-device placement fix from #3706 is now included in PyPI 1.4.16, uploaded September 18. A source checkout is no longer required just to obtain this fix:
python -m pip install "funasr==1.4.16"I verified the actual published wheel against PyPI's SHA-256; its
funasr/auto/auto_model.pyis byte-for-byte identical to the #3706 merged implementation. The matching English and Chinese API version-note correction is now merged in PR #3724, after all PR CI checks passed. Please still follow the installation guide for dependencies and record the importedfunasr.__file__when checking an existing environment.This supersedes only the release-availability caveat in my September 16 comment. It does not establish the original M2/121-minute performance result, which has not been independently retested; keeping this issue open for that confirmation.
Summary
AutoModelforcesvad_kwargs["device"]to equal the main ASR model's device (funasr/auto/auto_model.py:470). There is no way to run VAD on CPU while the ASR model runs on GPU/MPS. On Apple Silicon this is a measurable ~5x performance regression for the VAD stage, because the FSMN streaming VAD emits many tiny per-frame forwards that suffer from MPS per-op kernel-launch overhead.Environment
torch.backends.mps.is_available() = TrueMeasured impact
Same 121-minute (7285s) audio, FSMN VAD only:
cpumpsVAD on MPS is 5.2x slower than CPU. Full pipeline (
paraformer-large+fsmn-vad+ct-punc), same 121-min audio:device='mps'(VAD+ASR both on MPS)Mixed device saves ~40% end-to-end. ASR (paraformer) genuinely benefits from MPS (large batched matmuls); VAD does not.
Root cause
funasr/auto/auto_model.py(1.4.14):So even passing
vad_kwargs={"device": "cpu"}is overwritten.inference_with_vad()then runsself.inference(model=self.vad_model, kwargs=self.vad_kwargs)withvad_kwargs["device"]fixed to the main device, so feature tensors land on the main device and VAD weights must match.Feature request
Expose a way to place the VAD model on a device independent of the ASR model, e.g.:
which would build the VAD on
vad_deviceand move feature tensors fed toself.vad_modelontovad_devicebefore VADforward, while the ASR model stays ondevice. (punc/spk sub-models have the same hardcoded coupling at lines 483/499, so a general per-submodel device would be ideal.)This matters most on Apple Silicon today, but the pattern (streaming VAD = many tiny forwards) is device-agnostic: any accelerator with high per-op launch overhead is hurt by forcing VAD onto it.
Minimal reproduction
Workaround (1.4.14)
Build the main model on the accelerator, then replace
MODEL.vad_modelwith a separately-built CPU VAD instance and patchComputeScoresto move feature tensors onto the CPU device:Works (verified, 156s vs 259s) but fragile across versions, hence this request.
Happy to turn this into a PR — the design question is whether to honor
vad_kwargs["device"]when explicitly provided (smallest change) vs. add a dedicatedvad_deviceargument.