Skip to content

[Feature Request] Allow VAD to run on a different device than the ASR model (Apple Silicon MPS regression: VAD 5x slower than CPU) #3701

Description

@damnwenxi

Summary

AutoModel forces vad_kwargs["device"] to equal the main ASR model's device (funasr/auto/auto_model.py:470). There is no way to run VAD on CPU while the ASR model runs on GPU/MPS. On Apple Silicon this is a measurable ~5x performance regression for the VAD stage, because the FSMN streaming VAD emits many tiny per-frame forwards that suffer from MPS per-op kernel-launch overhead.

Environment

  • macOS / Apple M2 (8-core, 16GB)
  • funasr 1.4.14
  • torch 2.14.0, torch.backends.mps.is_available() = True

Measured impact

Same 121-minute (7285s) audio, FSMN VAD only:

Device VAD wall-clock Realtime factor
cpu 23.6s 308x
mps 123.0s 59x

VAD on MPS is 5.2x slower than CPU. Full pipeline (paraformer-large + fsmn-vad + ct-punc), same 121-min audio:

Config Total Notes
device='mps' (VAD+ASR both on MPS) ~259s VAD=123s, ASR=91s
VAD on CPU + ASR on MPS (mixed) ~156s VAD=24s, ASR=91s

Mixed device saves ~40% end-to-end. ASR (paraformer) genuinely benefits from MPS (large batched matmuls); VAD does not.

Root cause

funasr/auto/auto_model.py (1.4.14):

# AutoModel.__init__, ~line 465-470
vad_kwargs = {} if kwargs.get("vad_kwargs", {}) is None else kwargs.get("vad_kwargs", {})
if vad_model is not None:
    vad_kwargs["model"] = vad_model
    vad_kwargs["model_revision"] = kwargs.get("vad_model_revision", "master")
    vad_kwargs["device"] = kwargs["device"]   # <-- hardcoded to main device

So even passing vad_kwargs={"device": "cpu"} is overwritten. inference_with_vad() then runs self.inference(model=self.vad_model, kwargs=self.vad_kwargs) with vad_kwargs["device"] fixed to the main device, so feature tensors land on the main device and VAD weights must match.

Feature request

Expose a way to place the VAD model on a device independent of the ASR model, e.g.:

AutoModel(model=..., vad_model=..., punc_model=..., device='mps', vad_device='cpu')

which would build the VAD on vad_device and move feature tensors fed to self.vad_model onto vad_device before VAD forward, while the ASR model stays on device. (punc/spk sub-models have the same hardcoded coupling at lines 483/499, so a general per-submodel device would be ideal.)

This matters most on Apple Silicon today, but the pattern (streaming VAD = many tiny forwards) is device-agnostic: any accelerator with high per-op launch overhead is hurt by forcing VAD onto it.

Minimal reproduction

import time
from pathlib import Path
from funasr import AutoModel

models = [Path.home()/'.cache/modelscope/models'/('iic--'+n)/'snapshots/master' for n in [
    'speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch',
    'speech_fsmn_vad_zh-cn-16k-common-pytorch',
    'punc_ct-transformer_cn-en-common-vocab471067-large',
]]
for dev in ('cpu', 'mps'):
    m = AutoModel(model=str(models[1]), device=dev, disable_update=True, disable_pbar=True)
    t = time.monotonic()
    m.generate(input='your_16k_mono.wav', max_single_segment_time=60000)
    print(dev, f'{time.monotonic()-t:.1f}s')

Workaround (1.4.14)

Build the main model on the accelerator, then replace MODEL.vad_model with a separately-built CPU VAD instance and patch ComputeScores to move feature tensors onto the CPU device:

main = AutoModel(model=asr, vad_model=vad, punc_model=punc, device='mps', ...)
cpu_vad = AutoModel(model=vad, device='cpu', ...).model
_orig = cpu_vad.ComputeScores
cpu_vad.ComputeScores = lambda feats, cache=None: _orig(feats.to('cpu') if hasattr(feats,'to') and feats.device.type!='cpu' else feats, cache=cache)
main.vad_model = cpu_vad
main.vad_kwargs['device'] = 'cpu'

Works (verified, 156s vs 259s) but fragile across versions, hence this request.

Happy to turn this into a PR — the design question is whether to honor vad_kwargs["device"] when explicitly provided (smallest change) vs. add a dedicated vad_device argument.

Activity

  1. LauraGPT commented on Sep 12, 2026

    @LauraGPT
    Collaborator

    Thanks for the detailed report and the offer to contribute a PR. I recommend honoring the existing nested device options rather than adding a second family of vad_device/punc_device/spk_device arguments. The proposed interface is:

    # Proposed behavior, not supported correctly by current main yet.
    model = AutoModel(
        model=asr,
        device="mps",
        vad_model=vad,
        vad_kwargs={"device": "cpu"},
        punc_model=punc,
        punc_kwargs={"device": "cpu"},
    )

    I checked current main 486b4b7ceb27b72db24a84c6e1be7cf6c5be6609. A CPU-only configuration probe using the actual AutoModel orchestration and stand-in models confirmed three relevant paths:

    • Construction overwrites an explicit device="cpu" in all three submodel dictionaries. It also modifies the caller's dictionaries in place.
    • Even after establishing a CPU VAD baseline, a shared generate(device="cuda:0") config changes the device passed to VAD back to cuda:0; the next call without that override resets to CPU. This is the deep_update path in both inference_with_vad and inference, so changing only the constructor assignments would not cover it.
    • The speaker-embedding call at auto_model.py:982 receives the ASR kwargs, not self.spk_kwargs. In the probe it received the ASR device and lost a speaker-only sentinel setting, despite a CPU speaker baseline.

    For a PR, please make the contract consistent across VAD, punctuation and speaker models:

    1. An explicitly supplied submodel device takes precedence; when omitted, retain the existing inheritance from the resolved ASR device. Copy caller-owned config dictionaries before filling defaults.
    2. Treat the selected submodel device as model placement, not an ordinary shared per-call option. Preserve it through runtime config merging and repeated-call resets, keeping weights and input features on the same device. Do not imply that changing a runtime device string moves already-loaded weights; runtime migration would need a separate explicit contract.
    3. Pass the speaker model its own resolved config. Cover punctuation both with and without VAD, since both routes merge shared runtime options.
    4. Add regressions for explicit overrides, omitted-device inheritance, unchanged caller dictionaries, shared runtime options, two successive calls, and the actual per-model inference boundaries. Please also exercise unavailable-device fallback without claiming hardware execution from mocked availability.

    A PR covering this would be welcome. Please include an actual mixed-device smoke test where hardware is available, checking model/feature device agreement and successful output. Your original Apple M2 audio is particularly valuable for validating the reported performance after the fix. My probe used CPU tensors only (cuda:0 was a configuration label), so I have not independently reproduced the MPS timing figures or a hardware device-mismatch crash. No ComputeScores monkeypatch should be required by the final public API.

  2. LauraGPT commented on Sep 16, 2026

    @LauraGPT
    Collaborator

    Update: the nested-device fix has now been merged in PR #3706, at commit 4d8e087. The earlier comment's “not supported correctly by current main yet” note describes the old base and no longer applies to this merged source.

    The public configuration is documented in the merged English and Chinese API guides:

    • An explicit nested device in vad_kwargs, punc_kwargs, or spk_kwargs is preserved; omitted values inherit the resolved ASR device.
    • Shared per-call options do not migrate already-loaded models.
    • Caller-owned submodel dictionaries are copied, and the speaker model receives its own configuration.

    I checked that all six merged PR files match the final candidate. Compared with the previously reviewed eee63614, the only follow-up change is an optional-clustering test-fixture adjustment; the production AutoModel source is unchanged.

    This is a source merge, not a new PyPI release announcement. The original M2/121-minute performance case has not been independently retested, and the author's M4 measurements remain workload-specific. Keeping this issue open for original-environment confirmation; no internal VAD replacement or ComputeScores patch should be necessary for the documented placement configuration.

  3. LauraGPT commented on Sep 21, 2026

    @LauraGPT
    Collaborator

    Release availability update: the mixed-device placement fix from #3706 is now included in PyPI 1.4.16, uploaded September 18. A source checkout is no longer required just to obtain this fix:

    python -m pip install "funasr==1.4.16"

    I verified the actual published wheel against PyPI's SHA-256; its funasr/auto/auto_model.py is byte-for-byte identical to the #3706 merged implementation. The matching English and Chinese API version-note correction is now merged in PR #3724, after all PR CI checks passed. Please still follow the installation guide for dependencies and record the imported funasr.__file__ when checking an existing environment.

    This supersedes only the release-availability caveat in my September 16 comment. It does not establish the original M2/121-minute performance result, which has not been independently retested; keeping this issue open for that confirmation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions