Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 14 additions & 11 deletions fern/docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -338,12 +338,21 @@ navigation:
contents:
- page: Overview
path: gpt-live/overview.mdx
- page: Design your assistant
- page: Quickstart
path: gpt-live/quickstart.mdx
- page: Design conversations
path: gpt-live/design.mdx
- page: Configuration
- page: Build with tools
path: gpt-live/tools.mdx
- page: Migrate to GPT-Live
path: gpt-live/migrate.mdx
- page: Test and improve
path: gpt-live/testing.mdx
- page: Settings and compatibility
path: gpt-live/configuration.mdx
- page: Limitations and FAQs
path: gpt-live/limitations.mdx
hidden: true
- page: Custom keywords
path: customization/custom-keywords.mdx
icon: fa-light fa-bullseye
Expand Down Expand Up @@ -1114,22 +1123,16 @@ redirects:
destination: /providers/voice/overview
- source: /providers/gpt-live
destination: /gpt-live/overview
- source: /gpt-live/quickstart
destination: /gpt-live/configuration#create-an-assistant
- source: /gpt-live/voices
destination: /gpt-live/configuration#voices
- source: /gpt-live/prompts
destination: /gpt-live/design#speaker-and-reasoner-roles
destination: /gpt-live/design#turn-the-design-into-prompts
- source: /gpt-live/conversation-design
destination: /gpt-live/design
- source: /gpt-live/operations
destination: /gpt-live/configuration#review-and-monitor-calls
- source: /gpt-live/tools
destination: /gpt-live/configuration#connect-your-tools
- source: /gpt-live/testing
destination: /gpt-live/design#testing-conversations
destination: /gpt-live/testing#watch-production-calls
- source: /gpt-live/faqs
destination: /gpt-live/limitations
destination: /gpt-live/configuration#compatibility
- source: /test/test-suites
destination: /test/voice-testing
- source: /test/chat-testing
Expand Down
478 changes: 157 additions & 321 deletions fern/gpt-live/configuration.mdx

Large diffs are not rendered by default.

653 changes: 485 additions & 168 deletions fern/gpt-live/design.mdx

Large diffs are not rendered by default.

127 changes: 15 additions & 112 deletions fern/gpt-live/limitations.mdx
Original file line number Diff line number Diff line change
@@ -1,153 +1,56 @@
---
title: GPT-Live limitations and FAQs
subtitle: Check compatibility before you build or migrate
description: Check GPT-Live support for squads, call control, commentary injection, transfers, simulations, voices, fallbacks, recordings, and post-call data in Vapi.
subtitle: This page has moved
description: The GPT-Live limitations and FAQs now live in Settings and compatibility and the other GPT-Live guides. Each former section links to its new home.
slug: gpt-live/limitations
---

GPT-Live has a different set of supported features from Vapi's other voice architectures. Use this page to check the requirements of your application before migrating.

These limits describe **Vapi's GPT-Live integration**. A capability in OpenAI's direct API isn't necessarily exposed through Vapi.

**Jump to:** [Calls and voices](#calls-and-voice) · [Transfers](#transfers-and-keypad-input) · [Squads](#can-i-use-an-existing-squad) · [Live control](#can-i-inject-commentary-or-control-speech-during-a-call) · [Idle messages](#can-i-use-idle-message-hooks) · [Testing](#can-i-run-simulations-and-evals) · [Recordings](#recordings-and-post-call-data)
The content of this page has moved. Supported features are in [Settings and compatibility](/gpt-live/configuration#compatibility), and the answers to these questions are in the guides linked below.

## Calls and voice

| Feature | Support |
| --- | --- |
| Browser WebRTC, native Twilio, and Vapi SIP | Supported |
| Raw WebSocket | Supported with mono, signed 16-bit little-endian PCM at 24 kHz. Set the rate explicitly |
| Native Telnyx or Vonage | Unsupported |
| 22 preset OpenAI voices | Supported. See [voice previews](/gpt-live/configuration#voices) |
| Custom or cloned voices, or other voice providers | Unsupported |
| Changing voice during a call | Unsupported. Save the assistant and start a new call |
| Model and voice fallbacks | Unsupported. GPT-Live also cannot be a fallback model |
| Reconnecting to an existing GPT-Live session | Unsupported. Start a new call and check the status of pending work |

See [Connect a call](/gpt-live/configuration#connect-a-call) for working connection examples.
See [Calls and voice](/gpt-live/configuration#calls-and-voice) in Settings and compatibility.

## Transfers and keypad input

Function tools, API request tools, MCP tools, and `endCall` are supported. MCP tools run through the reasoner; see [MCP configuration](/gpt-live/configuration#mcp-tools). The other supported tool types, `transferCall` and `dtmf`, depend on the call connection:

| Feature | Support |
| --- | --- |
| Blind transfer | Supported on native Twilio and Vapi SIP to phone or SIP destinations. Configure destinations for a `transferCall` tool, or supply one in an HTTP `transfer` request |
| Transfer on browser or raw WebSocket calls | Unsupported |
| Warm transfer or generated transfer summary | Unsupported |
| Transfer fallback plans or SIP `bye` verb | Unsupported |
| SIP transfer with extension dialing | Unsupported. Use a direct destination |
| Outgoing DTMF | Supported on SIP using RTP DTMF. Pass digits as a string, for example `"0011*#"` |
| Outgoing DTMF on Twilio, browser, or raw WebSocket calls | Unsupported |
| SIP INFO DTMF | Unsupported |
| Incoming keypad collection | Unsupported. Leave `keypadInputPlan.enabled` disabled |

Other tool types are unsupported. Expose external lookups through a function, API request, or MCP tool. Carrier acceptance of a transfer doesn't establish that the destination answered.
See [Transfers and keypad input](/gpt-live/configuration#transfers-and-keypad-input) in Settings and compatibility, and [Transfer to a person](/gpt-live/tools#transfer-to-a-person).

## Can I use an existing squad?

Squads and assistant handoffs are currently unsupported. GPT-Live's speaker and reasoner work within one assistant. Delegation doesn't switch the caller to another assistant.

Early testing suggests some use cases no longer need the same divisions between assistants. See [Start with one assistant](/gpt-live/design#start-with-one-assistant) for how to rethink the architecture while keeping your business requirements.
Squads and handoffs aren't supported with GPT-Live. See [Migrate to GPT-Live: From a squad](/gpt-live/migrate#from-a-squad).

## Can I inject commentary or control speech during a call?

Yes. Send HTTP requests to the call's `monitor.controlUrl` to:

- End the call with `end-call`.
- Append `commentary`, `thinking`, or `instructions` to the speaker with `append-context`.
- Initiate a cold transfer with `transfer` on native Twilio or Vapi SIP calls.

See [Live call control](/gpt-live/configuration#live-call-control) for request examples, authentication, and response behavior. Appended context affects the speaker; it doesn't update reasoner history or directly trigger tools. Submission doesn't guarantee exact wording or completed speech.

The raw WebSocket connection supports audio streaming only. Send controls over HTTP, not WebSocket text frames or SDK data channels. Classic HTTP `say`, `add-message`, and mute/unmute commands remain unsupported. The `say` and `message.add` actions in [idle-message hooks](#can-i-use-idle-message-hooks) are separate assistant configuration.

A caller can ask the assistant to slow down or explain differently, and you can prompt it to respond to those requests. That conversational behavior is separate from an application sending a control command.
See [Live call control](/gpt-live/configuration#live-call-control).

## Can I use idle-message hooks?

Yes. GPT-Live supports `customer.speech.timeout` hooks with `say.prompt`, `message.add`, and inline `endCall` actions. Idle messages are generated from the prompt; exact wording isn't guaranteed. Use separate hooks with staggered timeouts for multiple check-ins. Existing trigger limits and reset modes apply; a check-in doesn't postpone later hooks in the same silence window.

Caller speech cancels remaining check-ins for that window. A `say` action requests one response and instructs the speaker not to repeat or resume it after an interruption. A hook's explicit `endCall` is different: once it fires, caller speech doesn't cancel the pending hangup.

Saved tool references, function calls, transfers, and other actions aren't supported by these hooks. Support for idle messages doesn't imply support for every assistant hook event. See [Idle messages](/gpt-live/configuration#idle-messages) for examples and supported actions. Hook speech is generated and remains subject to the wording limitations below.
Yes. See [When the caller goes quiet](/gpt-live/design#when-the-caller-goes-quiet) for an example and [Idle messages](/gpt-live/configuration#idle-messages) for the settings.

## Can I guarantee exact speech or a fixed pause?

GPT-Live generates speech and may paraphrase supplied text. Speaker instructions and personality packs guide delivery. They don't provide exact playback, a fixed speaking rate, or a programmatic pause while work completes.

Audio URL greetings and the generated-message first-message mode are unsupported. Use a text `firstMessage` with `assistant-speaks-first`, or let the assistant wait for the caller. Without usable greeting text, it waits.

If exact prerecorded wording or strict control of every spoken step is essential, choose an architecture with those controls. See [voice design](/gpt-live/design#speaking-style) for the choices available through prompting.
No. GPT-Live generates all of its speech. See [Speaking style](/gpt-live/design#speaking-style).

## Does the reasoner get the conversation transcript?

Yes. Each delegation receives the available caller and assistant transcript up to that point, along with the reasoner's instructions, tool definitions, and retained reasoner/tool history. It isn't limited to a short snippet chosen by the speaker.

It doesn't receive raw audio or automatically inherit the speaker prompt. The transcript is a snapshot at delegation time; later speech doesn't automatically update a delegation already in progress. See [reasoner context](/gpt-live/design#what-context-does-the-reasoner-receive) for the context comparison and correction example, and [prompt guidance](/gpt-live/design#where-should-an-instruction-go) for what to put in each prompt.
See [What context does the reasoner receive?](/gpt-live/design#what-context-does-the-reasoner-receive)

## Which existing settings no longer apply?

Classic transcriber selection, text-to-speech controls, and start/stop-speaking plans don't configure the GPT-Live conversation. General model fields such as `temperature` and `maxTokens` don't tune this integration.

Use the speaker prompt for conversation behavior and the reasoner settings for task reasoning. The reasoner currently supports only the OpenAI models listed in [Model settings](/gpt-live/configuration#model-settings).
See [Settings that no longer apply](/gpt-live/configuration#settings-that-no-longer-apply) and the [settings map for migration](/gpt-live/migrate#4-map-the-settings).

## Does an interruption cancel a tool call?

No. Interrupting speech changes the conversation, not an action already submitted to your service. Your application needs to check whether that action is pending, completed, or cancellable.

If the caller corrects a date during a lookup, use the result that matches the updated request. For an example, see [When callers change the request](/gpt-live/design#when-callers-change-the-request).
No. See [When callers change the request](/gpt-live/design#when-callers-change-the-request).

## Can I run Simulations and Evals?

Yes. **Voice Simulations support GPT-Live** as the assistant under test, the AI tester, or both. You can also pair a GPT-Live participant with a classic voice assistant.

| Simulation setup | Support |
| --- | --- |
| GPT-Live target with a classic AI tester | Supported in Voice mode |
| GPT-Live AI tester with a classic assistant target | Supported in Voice mode |
| GPT-Live tester and target | Supported in Voice mode |
| Chat mode with either participant using GPT-Live | Unsupported; choose Voice |
| Squad runs with a GPT-Live participant | Unsupported; choose a single assistant target |

To test your assistant:

1. Follow the [Simulations quickstart](/observability/simulations-quickstart), selecting your GPT-Live assistant as the target.
2. Configure the scenario, AI tester personality, and success criteria. If you use a GPT-Live tester, select a supported native OpenAI voice.
3. Run the suite in **Voice** mode, then review its transcript, evaluations, and available recording.

Existing GPT-Live organization access requirements still apply. Simulations handle the participants' audio formats, including mixed GPT-Live and classic runs; you don't need to build an audio bridge yourself. [Tool mocks](/observability/simulations-advanced#mock-tool-responses) are supported for GPT-Live simulation tools. Unmocked tools can perform real actions, so configure mocks or sandbox services before running scenarios that change external data.

Voice Simulations let you exercise speech, timing, interruptions, and tool behavior. Listen to the recording as well as checking evaluations: a passing criterion doesn't establish every aspect of conversation quality. [Evals](/observability/evals-quickstart) remain useful for supported model decisions, but don't replace testing the voice interaction.
See [Repeat scenarios with Voice Simulations](/gpt-live/testing#repeat-scenarios-with-voice-simulations).

## Recordings and post-call data

GPT-Live supports post-call analysis, structured outputs, scorecards, Boards, and Monitoring. Availability depends on your configuration and retained call data:

| Setting or feature | Behavior |
| --- | --- |
| Recording disabled in the assistant | No recording is created |
| Recording-consent plan configured | Recording is disabled because GPT-Live doesn't collect that consent |
| Transcript-based analysis | Requires the relevant conversation data |
| Audio-based analysis | Requires an available recording |
| Monitoring with Zero Data Retention | Unsupported |
| Live listen or monitor sockets | Unsupported. Post-call Monitoring doesn't provide a live audio connection |

See [Review and monitor calls](/gpt-live/configuration#review-and-monitor-calls) for setup and result fields.
See [Recordings and post-call data](/gpt-live/configuration#recordings-and-post-call-data).

## Troubleshooting

Use the symptom to find the setting or handler to check:

| Symptom | Check | Fix |
| --- | --- | --- |
| GPT-Live is unavailable | Organization feature access | Contact Vapi to enable access. |
| Voice unsupported | `voice.provider` and `voice.voiceId` | Use OpenAI and a [supported voice ID](/gpt-live/configuration#voices). |
| Tool configuration unsupported | Tool type and resolved names | Use supported types and unique names. Inspect saved tools too. |
| Call fails on WebSocket startup | Audio format and sample rate | Set raw `pcm_s16le` at `24000` Hz explicitly. |
| WebSocket audio sounds distorted | Sample encoding, rate, and channels | Convert the source to mono, signed 16-bit little-endian PCM at 24 kHz. |
| Assistant never greets | First-message mode and usable text | Set `assistant-speaks-first` and a text `firstMessage`. |
| Assistant promises a lookup but no tool runs | Both prompts and attached tools | Add a concrete delegation rule and make the tool available to the reasoner. |
| Tool result never reaches the conversation | Handler response and matching ID | Return the actual `toolCallId`. Return the final async result through the original handler request. |
| Transfer fails | Transport and configured destination | Use a supported blind transfer. Check carrier requirements and destination restrictions. |
| No recording | Recording settings and consent plan | Check the documented recording behavior before the test. |
See [Troubleshooting](/gpt-live/testing#troubleshooting) in Test and improve.
Loading
Loading