Complete guide to using Text-to-Speech (TTS) and Speech-to-Text (STT) with AgentSea ADK.
- Overview
- Speech-to-Text (STT)
- Text-to-Speech (TTS)
- Voice Agent
- Supported Providers
- Examples
- Best Practices
AgentSea ADK includes comprehensive voice support for building voice-enabled AI agents:
- Speech-to-Text (STT) - Transcribe audio to text
- Text-to-Speech (TTS) - Synthesize speech from text
- Voice Agent - Wrapper that combines both for voice conversations
- Multiple Providers - Cloud and local options
- Streaming - Real-time audio streaming
- Multiple Languages - Support for many languages
High-quality transcription with OpenAI's Whisper model.
import { OpenAIWhisperProvider } from '@lov3kaizen/agentsea-core';
const sttProvider = new OpenAIWhisperProvider(process.env.OPENAI_API_KEY);
// Transcribe audio file
const result = await sttProvider.transcribe('./audio.mp3', {
model: 'whisper-1',
language: 'en',
responseFormat: 'verbose_json',
});
console.log('Text:', result.text);
console.log('Language:', result.language);
console.log('Duration:', result.duration);
// Access segments with timestamps
result.segments?.forEach((segment) => {
console.log(`[${segment.start}s - ${segment.end}s]: ${segment.text}`);
});
// Access word-level timestamps
result.words?.forEach((word) => {
console.log(`${word.word} at ${word.start}s`);
});Features:
- ✅ High accuracy
- ✅ 99+ languages
- ✅ Timestamps (segment and word-level)
- ✅ Speaker diarization
- ❌ No streaming
Supported Languages: English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Chinese, Japanese, Korean, Arabic, Hindi, and 90+ more.
Cost-effective OpenAI-compatible transcription using Whisper v3.
import { LemonFoxSTTProvider } from '@lov3kaizen/agentsea-core';
const sttProvider = new LemonFoxSTTProvider(process.env.LEMONFOX_API_KEY);
// Transcribe audio file
const result = await sttProvider.transcribe('./audio.mp3', {
model: 'whisper-1',
language: 'en',
responseFormat: 'verbose_json',
});
console.log('Text:', result.text);
console.log('Language:', result.language);
console.log('Duration:', result.duration);Alternative: Using OpenAI Whisper with custom baseURL:
import { OpenAIWhisperProvider } from '@lov3kaizen/agentsea-core';
const sttProvider = new OpenAIWhisperProvider({
apiKey: process.env.LEMONFOX_API_KEY,
baseURL: 'https://api.lemonfox.ai/v1',
});Features:
- ✅ High accuracy (Whisper v3)
- ✅ 100+ languages
- ✅ Timestamps (segment and word-level)
- ✅ Speaker diarization
- ✅ OpenAI-compatible API
- ✅ Cost-effective ($0.50 per 3 hours)
- ❌ No streaming
Run Whisper locally for complete privacy.
import { LocalWhisperProvider } from '@lov3kaizen/agentsea-core';
const sttProvider = new LocalWhisperProvider({
whisperPath: '/path/to/whisper',
modelPath: '/path/to/ggml-base.bin',
});
// Check if installed
if (!(await sttProvider.isInstalled())) {
console.log(sttProvider.getInstallInstructions());
return;
}
const result = await sttProvider.transcribe('./audio.wav', {
model: 'base',
language: 'en',
});Installation:
Option 1: whisper.cpp (recommended)
git clone https://github.com/ggerganov/whisper.cpp
cd whisper.cpp
make
./models/download-ggml-model.sh baseOption 2: Python Whisper
pip install openai-whisperFeatures:
- ✅ Complete privacy
- ✅ No API costs
- ✅ Offline capability
- ❌ No streaming
- ❌ Slower than cloud
High-quality voices with OpenAI's TTS models.
import { OpenAITTSProvider } from '@lov3kaizen/agentsea-core';
const ttsProvider = new OpenAITTSProvider(process.env.OPENAI_API_KEY);
// Synthesize speech
const result = await ttsProvider.synthesize('Hello, world!', {
model: 'tts-1-hd',
voice: 'nova',
speed: 1.0,
format: 'mp3',
});
// Save audio
import { writeFileSync } from 'fs';
writeFileSync('./output.mp3', result.audio);
// Stream for faster response
for await (const chunk of ttsProvider.synthesizeStream('Long text...', {
voice: 'alloy',
})) {
// Process audio chunks
}Available Voices:
alloy- Neutral voiceecho- Male voicefable- Neutral voiceonyx- Male voicenova- Female voiceshimmer- Female voice
Models:
tts-1- Faster, lower latencytts-1-hd- Higher quality
Features:
- ✅ High quality
- ✅ Multiple voices
- ✅ Streaming support
- ✅ Speed control
- ✅ Multiple formats (mp3, opus, aac, flac, wav, pcm)
Cost-effective OpenAI-compatible text-to-speech with 50+ voices.
import { LemonFoxTTSProvider } from '@lov3kaizen/agentsea-core';
const ttsProvider = new LemonFoxTTSProvider(process.env.LEMONFOX_API_KEY);
// Synthesize speech
const result = await ttsProvider.synthesize('Hello, world!', {
model: 'tts-1',
voice: 'sarah', // or any OpenAI-compatible voice like 'nova'
format: 'mp3',
});
// Save audio
import { writeFileSync } from 'fs';
writeFileSync('./output.mp3', result.audio);
// Stream for faster response
for await (const chunk of ttsProvider.synthesizeStream('Long text...', {
voice: 'alloy',
})) {
// Process audio chunks
}Alternative: Using OpenAI TTS with custom baseURL:
import { OpenAITTSProvider } from '@lov3kaizen/agentsea-core';
const ttsProvider = new OpenAITTSProvider({
apiKey: process.env.LEMONFOX_API_KEY,
baseURL: 'https://api.lemonfox.ai/v1',
});Features:
- ✅ 50+ voices across 8 languages
- ✅ OpenAI-compatible API
- ✅ Streaming support
- ✅ Multiple formats (mp3, opus, aac, flac, wav, pcm)
- ✅ Low latency
- ✅ Cost-effective ($2.50 per 1M characters, up to 90% savings)
Premium voice synthesis with voice cloning.
import { ElevenLabsTTSProvider } from '@lov3kaizen/agentsea-core';
const ttsProvider = new ElevenLabsTTSProvider({
apiKey: process.env.ELEVENLABS_API_KEY,
});
// List available voices
const voices = await ttsProvider.getVoices();
console.log('Available voices:', voices);
// Synthesize with specific voice
const result = await ttsProvider.synthesize('Hello!', {
voice: 'EXAVITQu4vr4xnSDxMaL', // Bella
model: 'eleven_multilingual_v2',
});
// Stream
for await (const chunk of ttsProvider.synthesizeStream('Text...')) {
// Process chunks
}Features:
- ✅ Highest quality
- ✅ Voice cloning
- ✅ Emotional range
- ✅ Multiple languages
- ✅ Streaming support
- 💰 Premium pricing
Fast, local neural TTS.
import { PiperTTSProvider } from '@lov3kaizen/agentsea-core';
const ttsProvider = new PiperTTSProvider({
piperPath: '/path/to/piper',
modelPath: '/path/to/en_US-lessac-medium.onnx',
});
// Check installation
if (!(await ttsProvider.isInstalled())) {
console.log(ttsProvider.getInstallInstructions());
return;
}
const result = await ttsProvider.synthesize('Hello!');
writeFileSync('./output.wav', result.audio);Installation:
# Download Piper
wget https://github.com/rhasspy/piper/releases/latest/download/piper_linux_x86_64.tar.gz
tar -xzf piper_linux_x86_64.tar.gz
# Download a voice model
wget https://huggingface.co/rhasspy/piper-voices/resolve/main/en/en_US/lessac/medium/en_US-lessac-medium.onnx
wget https://huggingface.co/rhasspy/piper-voices/resolve/main/en/en_US/lessac/medium/en_US-lessac-medium.onnx.jsonFeatures:
- ✅ Fast synthesis
- ✅ Complete privacy
- ✅ No API costs
- ✅ Multiple voices
- ❌ No streaming
- ❌ Lower quality than cloud
Combine STT and TTS for voice conversations.
import {
Agent,
AnthropicProvider,
ToolRegistry,
VoiceAgent,
OpenAIWhisperProvider,
OpenAITTSProvider,
} from '@lov3kaizen/agentsea-core';
// Create base agent
const provider = new AnthropicProvider(process.env.ANTHROPIC_API_KEY);
const toolRegistry = new ToolRegistry();
const agent = new Agent(
{
name: 'voice-assistant',
model: 'claude-opus-4-8',
provider: 'anthropic',
systemPrompt: 'You are a helpful voice assistant.',
description: 'Voice assistant',
},
provider,
toolRegistry,
);
// Create voice providers
const sttProvider = new OpenAIWhisperProvider(process.env.OPENAI_API_KEY);
const ttsProvider = new OpenAITTSProvider(process.env.OPENAI_API_KEY);
// Create voice agent
const voiceAgent = new VoiceAgent(agent, {
sttProvider,
ttsProvider,
ttsConfig: {
voice: 'nova',
model: 'tts-1',
},
autoSpeak: true, // Automatically synthesize responses
});
// Process voice input
const audioInput = readFileSync('./user-audio.mp3');
const result = await voiceAgent.processVoice(audioInput, context);
console.log('User said:', result.text);
console.log('Assistant response:', result.response.content);
// Save audio response
writeFileSync('./response.mp3', result.audio!);// Multi-turn conversation
const context = {
conversationId: 'conv-1',
sessionData: {},
history: [],
};
// Turn 1
let result = await voiceAgent.processVoice(
readFileSync('./turn1.mp3'),
context,
);
console.log('Turn 1 - User:', result.text);
console.log('Turn 1 - Assistant:', result.response.content);
// Turn 2
result = await voiceAgent.processVoice(readFileSync('./turn2.mp3'), context);
console.log('Turn 2 - User:', result.text);
console.log('Turn 2 - Assistant:', result.response.content);
// Export full conversation
await voiceAgent.exportConversation('./conversation-export');// Get spoken response from text input
const result = await voiceAgent.speak('Tell me a joke', context);
console.log('Response:', result.text);
writeFileSync('./joke.mp3', result.audio);// Transcribe without agent processing
const text = await voiceAgent.transcribe(audioBuffer);
console.log('Transcription:', text);// Stream long responses
for await (const chunk of voiceAgent.synthesizeStream('Long text...')) {
// Play audio chunk immediately
}// Update TTS settings
voiceAgent.setTTSConfig({
voice: 'onyx',
speed: 1.2,
model: 'tts-1-hd',
});
// Update STT settings
voiceAgent.setSTTConfig({
language: 'es',
temperature: 0,
});
// Toggle auto-speak
voiceAgent.setAutoSpeak(false);| Provider | Quality | Speed | Cost | Privacy | Streaming | Languages |
|---|---|---|---|---|---|---|
| OpenAI Whisper | Excellent | Fast | $$ | Cloud | ❌ | 99+ |
| LemonFox STT | Excellent | Fast | $ (cheapest) | Cloud | ❌ | 100+ |
| Local Whisper | Excellent | Slow | Free | 100% | ❌ | 99+ |
| Provider | Quality | Speed | Cost | Privacy | Streaming | Voices |
|---|---|---|---|---|---|---|
| OpenAI TTS | Excellent | Fast | $ | Cloud | ✅ | 6 |
| LemonFox TTS | Excellent | Fast | $ (cheapest) | Cloud | ✅ | 50+ |
| ElevenLabs | Premium | Fast | $$$ | Cloud | ✅ | 100+ |
| Piper TTS | Good | Fast | Free | 100% | ❌ | 50+ |
const agent = new Agent(
{
name: 'customer-service',
model: 'claude-opus-4-8',
provider: 'anthropic',
systemPrompt: 'You are a helpful customer service representative.',
description: 'Customer service bot',
},
provider,
toolRegistry,
);
const voiceAgent = new VoiceAgent(agent, {
sttProvider: new OpenAIWhisperProvider(),
ttsProvider: new OpenAITTSProvider(),
ttsConfig: { voice: 'nova' }, // Friendly female voice
});
// Handle customer calls
const result = await voiceAgent.processVoice(customerAudio, context);const agent = new Agent(
{
name: 'language-tutor',
model: 'claude-opus-4-8',
provider: 'anthropic',
systemPrompt: 'You are a Spanish language tutor.',
description: 'Language tutor',
},
provider,
toolRegistry,
);
const voiceAgent = new VoiceAgent(agent, {
sttProvider: new OpenAIWhisperProvider(),
ttsProvider: new OpenAITTSProvider(),
sttConfig: { language: 'es' },
ttsConfig: { voice: 'alloy' },
});
// Practice conversations
const result = await voiceAgent.processVoice(studentAudio, context);const agent = new Agent(
{
name: 'podcast-host',
model: 'claude-opus-4-8',
provider: 'anthropic',
systemPrompt: 'You are an engaging podcast host.',
description: 'Podcast host',
},
provider,
toolRegistry,
);
const voiceAgent = new VoiceAgent(agent, {
sttProvider: new OpenAIWhisperProvider(),
ttsProvider: new ElevenLabsTTSProvider(),
ttsConfig: {
voice: 'professional-voice-id',
model: 'eleven_multilingual_v2',
},
});
// Generate podcast episode
const script = "Today we're discussing...";
const result = await voiceAgent.speak(script, context);
writeFileSync('./podcast-episode.mp3', result.audio);For Production:
- STT: OpenAI Whisper (accuracy + speed)
- TTS: OpenAI TTS or ElevenLabs (quality)
For Cost-Effective Production:
- STT: LemonFox ($0.50 per 3 hours - lowest on market)
- TTS: LemonFox ($2.50 per 1M chars - up to 90% savings)
For Development:
- STT: Local Whisper (no costs)
- TTS: Piper TTS (no costs)
For Privacy:
- STT: Local Whisper
- TTS: Piper TTS
// Use streaming for faster perceived response
for await (const chunk of voiceAgent.synthesizeStream(longText)) {
playAudio(chunk); // Start playing immediately
}
// Use faster models
const ttsProvider = new OpenAITTSProvider();
voiceAgent.setTTSConfig({
model: 'tts-1', // Faster than tts-1-hd
});// Split long transcriptions
const result = await sttProvider.transcribe(longAudio, {
responseFormat: 'verbose_json',
});
// Process segments individually
for (const segment of result.segments) {
console.log(`[${segment.start}s]: ${segment.text}`);
}try {
const result = await voiceAgent.processVoice(audio, context);
} catch (error) {
if (error.message.includes('audio format')) {
// Handle unsupported format
} else if (error.message.includes('rate limit')) {
// Handle rate limiting
} else {
// Handle other errors
}
}// List available voices
const voices = await ttsProvider.getVoices();
// Choose based on use case
const customerService = 'nova'; // Friendly female
const news = 'onyx'; // Professional male
const storytelling = 'fable'; // Expressive neutral// Cache common responses
const cache = new Map<string, Buffer>();
async function getCachedSpeech(text: string): Promise<Buffer> {
if (cache.has(text)) {
return cache.get(text)!;
}
const audio = await voiceAgent.synthesize(text);
cache.set(text, audio);
return audio;
}
// Pre-generate common phrases
const greetings = await Promise.all([
voiceAgent.synthesize('Hello!'),
voiceAgent.synthesize('How can I help you?'),
voiceAgent.synthesize('Thank you!'),
]);Set environment variables:
export OPENAI_API_KEY=your_key
export ELEVENLABS_API_KEY=your_key
export LEMONFOX_API_KEY=your_keyConvert audio to supported format:
ffmpeg -i input.m4a -ar 16000 output.wavInstall whisper.cpp or Python whisper and provide path:
const provider = new LocalWhisperProvider({
whisperPath: '/usr/local/bin/whisper',
});- Use smaller model:
baseinstead oflarge - Use OpenAI Whisper instead of local
- Reduce audio quality
- Use higher quality model:
tts-1-hd - Use ElevenLabs for premium quality
- Adjust speed:
speed: 0.9for clearer speech