A Flutter plugin for cross-platform real-time audio streaming. Provides low-latency audio input/output with simple Stream-based API for audio processing, recording, and visualization.
Try the PCM16 streaming + Gemini Live example in your browser: wamf.github.io/audio_io — paste your own Gemini API key, allow the microphone, and talk.
- Real-time audio streaming from microphone to Flutter
- Audio output/playback through speakers
- Cross-platform support (iOS, macOS, Android, Web, Linux, Windows)
- Simple Stream-based API
- PCM16 byte streams at 16/24/48 kHz for realtime voice APIs (e.g. Gemini Live)
- System-audio (loopback) capture on Windows, macOS, and the web — record what the machine is playing, alone or mixed with the microphone
- Configurable audio latency modes
- Optional dedicated audio isolate on FFI platforms
- Consistent data format across all platforms (Float64, 48kHz, mono)
- Low-latency audio processing
- Volume level monitoring
| Platform | Status | Implementation |
|---|---|---|
| iOS | ✅ Supported | Native (AVAudioEngine) |
| macOS | ✅ Supported | Native (AVAudioEngine) |
| Android | ✅ Supported | FFI (miniaudio) |
| Web | ✅ Supported | Web Audio API |
| Linux | ✅ Supported | FFI (miniaudio) |
| Windows | ✅ Supported | FFI (miniaudio) |
Add audio_io to your pubspec.yaml:
dependencies:
audio_io: ^0.6.1Add microphone usage description to your Info.plist:
<key>NSMicrophoneUsageDescription</key>
<string>This app needs access to the microphone for audio processing.</string>- Add microphone usage description to your
macos/Runner/Info.plist:
<key>NSMicrophoneUsageDescription</key>
<string>This app needs access to the microphone for audio processing.</string>- Enable audio input in your entitlements files:
macos/Runner/DebugProfile.entitlementsmacos/Runner/Release.entitlements
<key>com.apple.security.device.audio-input</key>
<true/>Add microphone permission to your android/app/src/main/AndroidManifest.xml:
<uses-permission android:name="android.permission.RECORD_AUDIO" />Request the microphone permission at runtime (for example with the
permission_handler package) before calling start(). If permission has
not been granted, start() throws an AudioIoException whose
isPermissionDenied is true.
import 'package:audio_io/audio_io.dart';
// Get the audio instance
final audioIo = AudioIo.instance;
// Configure latency (optional)
await audioIo.requestLatency(AudioIoLatency.Balanced);
// Start audio processing
await audioIo.start();
// Listen to input audio stream
audioIo.input.listen((audioData) {
// Process audio data (List<double>)
print('Received ${audioData.length} samples');
// Calculate volume level (RMS)
final sum = audioData.fold<double>(
0.0, (sum, sample) => sum + sample * sample);
final rms = sqrt(sum / audioData.length);
});
// Send audio to output (echo example)
audioIo.input.listen((data) {
audioIo.output.add(data);
});
// Stop audio processing
await audioIo.stop();The plugin supports three latency modes:
enum AudioIoLatency {
Realtime, // Lowest latency (~1.5ms buffer)
Balanced, // Balanced latency/CPU (~3ms buffer)
Powersave, // Lower CPU usage (~6ms buffer)
}
// Set before starting audio
await audioIo.requestLatency(AudioIoLatency.Realtime);Realtime voice APIs typically speak little-endian PCM16 at specific sample
rates (Gemini Live expects 16 kHz in / 24 kHz out). startWith configures
byte streams in the rate and format the API expects, while the engine keeps
its internal 48 kHz contract — resampling and conversion are handled for
you, on the audio rendering thread where the platform supports it:
final audioIo = AudioIo.instance;
await audioIo.startWith(const AudioIoConfig(
format: AudioIoFormat.pcm16,
sampleRate: AudioIoSampleRate.rate16000,
));
// Microphone as PCM16 bytes at 16 kHz
audioIo.inputBytes.listen(api.sendAudio);
// Play PCM16 bytes (decode + resample handled internally)
api.audioResponses.listen(audioIo.outputBytes.add);
// Barge-in: discard audio queued for playback but not yet rendered
await audioIo.clearOutput();sampleRate is a shorthand that applies to both directions. When your API
uses a different rate per direction — e.g. OpenAI Realtime and Gemini Live
capture at 16 kHz but return audio at 24 kHz — set inputSampleRate and
outputSampleRate explicitly instead. Each direction is resampled to and
from the engine's 48 kHz contract independently; either field falls back to
sampleRate when omitted, so existing single-rate callers are unaffected:
await audioIo.startWith(const AudioIoConfig(
format: AudioIoFormat.pcm16,
inputSampleRate: AudioIoSampleRate.rate16000, // mic bytes at 16 kHz
outputSampleRate: AudioIoSampleRate.rate24000, // playback bytes at 24 kHz
));The output ring is sized from the frame duration by default. Set
outputBufferDuration to size it independently — cap it low so a barge-in
drops less stale audio, or raise it so a burst producer (e.g. Gemini Live
returning several seconds at once) is not dropped. It is expressed in
seconds of the 48 kHz playback contract; each back end enforces a small
safety floor, so smaller values are clamped up:
await audioIo.startWith(const AudioIoConfig(
outputBufferDuration: 5, // hold up to ~5 s of queued playback
));
// Or on a running session, before pushing a burst:
await audioIo.requestOutputBufferDuration(0.3); // cap ~300 ms for low latencySee example_gemini_live for a complete voice conversation app, or try it
in the browser at wamf.github.io/audio_io.
By default the audio transport runs on the main isolate, which suits most apps and every platform. On the FFI back ends (Android, Windows, Linux) and on iOS/macOS you can opt into a dedicated audio isolate so device polling and native buffer copies are unaffected by main-isolate jank (heavy widget builds, GC):
await audioIo.startWith(const AudioIoConfig(
threading: AudioIoThreading.audioIsolate,
));iOS and macOS reach the AVAudioEngine ring buffers over FFI (the engine
lifecycle stays on the method channel), so the data plane can run on the
audio isolate just like the FFI back ends. Web silently falls back to
main-isolate operation. The input / output streams still surface on the
main isolate in every mode, so listener callbacks run there; move heavy DSP
out of the listener if it competes with UI work.
Set inputSource to capture the machine's audio mix instead of the
microphone. On the web this is backed by
getDisplayMedia:
startWith triggers the browser's share picker (it must run from a user
gesture — a button tap is fine), and the audio track from the chosen tab or
screen is piped into the same 48 kHz mono graph the microphone uses. The
video track is required by the picker but is stopped immediately.
await audioIo.startWith(const AudioIoConfig(
inputSource: AudioIoInputSource.systemAudio,
));Platform reality — read before relying on it:
- Chromium only. Chrome and Edge deliver an audio track; Firefox and
Safari implement
getDisplayMediabut return no audio track. On those browsers the input stream emits anAudioIoExceptionwithisSystemAudioUnsupported == true— listen to the stream'sonError(or catch it) rather than assuming audio will arrive. - Tab audio works on every Chromium desktop platform when the user shares a tab — the right UX for browser-hosted meetings ("share the Meet tab").
- Full system audio (sharing the whole screen) works on Windows and ChromeOS always, and on macOS only since Chrome 141 on macOS 14.2+.
- The captured tab keeps playing out of the speakers by default, so a listening app does not silence the source it is capturing.
- When the user clicks Stop sharing in the browser UI, the capture track
ends and the
inputstream completes (onDone); output/playback keeps running.
On desktop AudioIoInputSource.systemAudio is backed by WASAPI loopback
(Windows) and Core Audio process taps (macOS); it throws the same
isSystemAudioUnsupported error on platforms/back ends that cannot provide
it. See the example/ app for a share-a-tab listening demo.
By default the input stream captures the microphone. Set
inputSource: AudioIoInputSource.systemAudio to instead capture the machine's
audio mix — what is currently playing out of the speakers (meetings, media,
other apps). The captured frames arrive on the same input / inputBytes
stream, downmixed to mono and resampled to the configured rate, so consumers
are unchanged.
await audioIo.startWith(const AudioIoConfig(
inputSource: AudioIoInputSource.systemAudio,
));| Platform | System audio | Mic + system audio | Mechanism |
|---|---|---|---|
| Windows | ✅ Supported (build 20348+) | ⛔ | WASAPI loopback (ma_device_type_loopback) |
| macOS | ✅ Supported (14.2+) | ✅ Supported (14.2+) | Core Audio process taps, mixed by AVAudioEngine |
| Web | ✅ Chromium | ⛔ | getDisplayMedia (see above) |
| Linux | ⛔ Not yet | ⛔ | PulseAudio/PipeWire monitor sources — planned |
| Android / iOS | ⛔ Not supported | ⛔ | — |
Microphone + system audio. AudioIoInputSource.microphoneAndSystemAudio
sums both into the one mono input stream — for a voice assistant that must
keep hearing its user while it listens to a meeting playing on the machine.
macOS only for now; other back ends throw isSystemAudioUnsupported.
System audio is captured through a
Core Audio process tap
(macOS 14.2+): a private aggregate device pairs the default output device
with a global tap that excludes this process, so audio the app plays through
the output stream is not fed back into the input. If the plugin cannot
resolve its own process for that exclusion list, startWith throws
isSystemAudioCaptureFailed instead of capturing everything. For
systemAudio the tap writes straight into the input ring the Dart side
drains; for microphoneAndSystemAudio it is rendered into the plugin's
AVAudioEngine mixer alongside the microphone. The captured audio keeps
playing out of the speakers.
- Info.plist: add
NSAudioCaptureUsageDescription— the first tap triggers a System Audio Recording permission prompt (its own TCC category, separate from the microphone). Sandboxed apps also needcom.apple.security.device.audio-input. Keep the deployment target at 14.4 or newer if you can: earlier targets land the prompt in a different TCC category with different copy. - Permission is prompted, not pre-checked: there is no public API to read
the audio-capture grant. If the user denies it the tap produces silence
rather than an error;
tccutil reset AudioCapture <bundle-id>resets it while testing. The prompt only fires for a signed app, and TCC attributes the request to the responsible process: an app launched by thefluttertool from a terminal is attributed to that terminal, which has noNSAudioCaptureUsageDescription, so the request is denied silently. Launch the app from Finder oropento be prompted as the app itself, or grant the terminal System Audio Recording Only in System Settings > Privacy & Security. - Microphone:
permission_handlerhas no macOS implementation, so the plugin requests microphone access itself (NSMicrophoneUsageDescription) when a source that includes the microphone starts and access is not yet determined. A system-audio-only session never touches the microphone. - Errors: macOS older than 14.2 throws
isSystemAudioUnsupported; a tap or aggregate-device failure throws anAudioIoExceptionwithisSystemAudioCaptureFailedand the failing call plusOSStatusin the message. - Device changes: the aggregate follows the output device that was the
default when capture started; the engine's configuration-change reset
rebuilds it, and every
stoptears it down so the next start binds to the current default. If the rebuild fails, the session ends and the failure is reported onAudioIo.sessionErrors; the nextstartrebuilds from scratch. - Mixed source and two clocks: for
microphoneAndSystemAudiothe tap ring is written on the output device's clock and drained on the input device's clock, and it is not rate-matched. Built-in speakers with the built-in microphone share one clock and are fine; a USB or Bluetooth microphone with the internal speakers drifts, and the ring zero-fills or drops a buffer every few seconds — an audible click in the mixed stream. - What is tested: the XCTest target (
example/macos/RunnerTests) covers the tap's pure helpers only — the interleaved and planar downmix and the aggregate-device description. The capture itself, and the own-process exclusion, are asserted byexample/integration_test/system_audio_test.darton a macOS host with the grant: another process's speech is heard, then the app's own tone is not, in the same session.
Own-process exclusion. On Windows the host process is excluded from the
loopback capture (wasapi.loopbackProcessID + loopbackProcessExclude), so an
app that plays TTS through the output stream while capturing system audio does
not hear itself. The output stream keeps working in this mode: because a
WASAPI loopback device is capture-only, a separate playback device is opened
alongside it.
Windows minimum: build 20348 (Windows 11 / Windows Server 2022).
Process-excluded loopback uses the WASAPI VAD\Process_Loopback activation
path, which only exists from build 20348. On older Windows (e.g. Windows 10
19045) the native device fails to initialise; rather than silently dropping the
own-process exclusion and re-capturing the app's own output, startWith throws
the same AudioIoException with isSystemAudioUnsupported == true as the
unsupported platforms below, so the microphone-fallback pattern covers this case
too.
No permission prompt is required for loopback capture on Windows.
Unsupported platforms throw an AudioIoException with
isSystemAudioUnsupported == true from startWith rather than crashing the
engine, so callers can fall back to the microphone:
try {
await audioIo.startWith(
const AudioIoConfig(inputSource: AudioIoInputSource.systemAudio),
);
} on AudioIoException catch (e) {
if (e.isSystemAudioUnsupported) {
await audioIo.startWith(const AudioIoConfig()); // microphone fallback
} else {
rethrow;
}
}All platforms use a consistent audio format:
- Sample Rate: 48kHz (may adapt to device capabilities)
- Channels: Mono (1 channel)
- Data Type: Double precision floats (Float64/double)
- Stream Format: Chunks of audio samples as
List<double> - Internal Processing: Platform-specific (Float32 on native platforms)
- Dart SDK: >=3.0.0 <4.0.0
- Flutter: >=3.10.0
- iOS: 13.0 or higher
- macOS: 10.15 or higher
See LICENSE file for details.