diff --git a/fern/docs.yml b/fern/docs.yml index 01ed8b8ba..c7ba1c4ea 100644 --- a/fern/docs.yml +++ b/fern/docs.yml @@ -338,12 +338,21 @@ navigation: contents: - page: Overview path: gpt-live/overview.mdx - - page: Design your assistant + - page: Quickstart + path: gpt-live/quickstart.mdx + - page: Design conversations path: gpt-live/design.mdx - - page: Configuration + - page: Build with tools + path: gpt-live/tools.mdx + - page: Migrate to GPT-Live + path: gpt-live/migrate.mdx + - page: Test and improve + path: gpt-live/testing.mdx + - page: Settings and compatibility path: gpt-live/configuration.mdx - page: Limitations and FAQs path: gpt-live/limitations.mdx + hidden: true - page: Custom keywords path: customization/custom-keywords.mdx icon: fa-light fa-bullseye @@ -1114,22 +1123,16 @@ redirects: destination: /providers/voice/overview - source: /providers/gpt-live destination: /gpt-live/overview - - source: /gpt-live/quickstart - destination: /gpt-live/configuration#create-an-assistant - source: /gpt-live/voices destination: /gpt-live/configuration#voices - source: /gpt-live/prompts - destination: /gpt-live/design#speaker-and-reasoner-roles + destination: /gpt-live/design#turn-the-design-into-prompts - source: /gpt-live/conversation-design destination: /gpt-live/design - source: /gpt-live/operations - destination: /gpt-live/configuration#review-and-monitor-calls - - source: /gpt-live/tools - destination: /gpt-live/configuration#connect-your-tools - - source: /gpt-live/testing - destination: /gpt-live/design#testing-conversations + destination: /gpt-live/testing#watch-production-calls - source: /gpt-live/faqs - destination: /gpt-live/limitations + destination: /gpt-live/configuration#compatibility - source: /test/test-suites destination: /test/voice-testing - source: /test/chat-testing diff --git a/fern/gpt-live/configuration.mdx b/fern/gpt-live/configuration.mdx index c7b2af880..73a60d33e 100644 --- a/fern/gpt-live/configuration.mdx +++ b/fern/gpt-live/configuration.mdx @@ -1,160 +1,85 @@ --- -title: Configure GPT-Live -subtitle: Create an assistant, choose its voice, and connect your tools -description: Configure a GPT-Live assistant in Vapi using the dashboard or API, with speaker and reasoner settings, voices, idle-message hooks, tools, call connections, and analysis. +title: GPT-Live settings and compatibility +subtitle: Supported features, fields, voices, hooks, call connections, and live controls +description: Reference for GPT-Live in Vapi. Feature compatibility, speaker and reasoner fields, personality packs, voices with previews, greetings, idle-message hooks, call connections, and live call control. slug: gpt-live/configuration --- -Create an assistant that checks appointment availability and reads back the results. You'll configure the speaker and reasoner, connect a lookup tool, and make a test call. +This page is the reference for GPT-Live's settings and supported features. For how to use them in a conversation, start with [Design conversations](/gpt-live/design). -For help deciding what to put in each prompt, read [Design your assistant](/gpt-live/design). For feature restrictions, use [Limitations and FAQs](/gpt-live/limitations). +**Jump to:** [Compatibility](#compatibility) · [Model settings](#model-settings) · [Personality](#personality-and-language) · [Voices](#voices) · [Greetings](#greetings-and-duration) · [Idle messages](#idle-messages) · [Calls](#connect-a-call) · [Live control](#live-call-control) -**Jump to:** [Create an assistant](#create-an-assistant) · [Model settings](#model-settings) · [Personality](#personality-and-language) · [Voices](#voices) · [Idle messages](#idle-messages) · [Tools](#connect-your-tools) · [Calls](#connect-a-call) · [Live control](#live-call-control) · [Analysis](#review-and-monitor-calls) +## Compatibility -## Before you start +GPT-Live supports a different set of features from Vapi's other voice architectures. These limits describe Vapi's GPT-Live integration. A capability in OpenAI's API or in ChatGPT isn't necessarily available through Vapi. -- Enable GPT-Live for your Vapi organization. -- Have a browser and microphone ready for a test call. -- For API requests, use a Vapi private API key on your server or local machine. -- If you bring an OpenAI key, it needs access to GPT-Live and the reasoner you select. -- Prepare an HTTPS availability endpoint using the [tool response contract](#connect-your-tools). +### Calls and voice -The example uses a placeholder endpoint. Replace it with your handler before testing a lookup. It returns availability only. Booking requires a separate tool. - -## Create an assistant - -Choose the dashboard or API. Both paths create an assistant with an availability lookup and an end-call tool. - - - - - -Open [Assistants](https://dashboard.vapi.ai/assistants), select **Create Assistant**, and choose **GPT-Live**. Choose Marin as the voice. - - -In **Speaker**, enter: - -```text title="Speaker instructions" -Help callers check appointment availability. Speak warmly and briefly. -Ask for an explicit date and time zone, then delegate the lookup. -If the caller changes the request, use the updated details. -Explain the returned slots. Booking isn't connected yet. -Delegate requests to end the call. -``` +| Feature | Support | +| --- | --- | +| Browser WebRTC, native Twilio, and Vapi SIP | Supported | +| Raw WebSocket | Supported with mono, signed 16-bit little-endian PCM at 24 kHz. Set the rate explicitly | +| Native Telnyx or Vonage | Not supported | +| 22 preset OpenAI voices | Supported. See [voices](#voices) | +| Custom or cloned voices, or other voice providers | Not supported | +| Changing voice during a call | Not supported. Save the assistant and start a new call | +| Model and voice fallbacks | Not supported. GPT-Live also can't be a fallback model | +| Reconnecting to an existing GPT-Live session | Not supported. Start a new call and check the status of pending work | + +### Transfers and keypad input + +| Feature | Support | +| --- | --- | +| Blind transfer | Supported on native Twilio and Vapi SIP, to phone or SIP destinations. Use a `transferCall` tool with configured destinations, or an HTTP [`transfer`](#cold-transfer) request | +| Transfer on browser or raw WebSocket calls | Not supported | +| Warm transfer or generated transfer summary | Not supported | +| Transfer fallback plans or SIP `bye` verb | Not supported | +| SIP transfer with extension dialing | Not supported. Use a direct destination | +| Outgoing DTMF | Supported on SIP using RTP DTMF. Pass digits as a string, for example `"0011*#"` | +| Outgoing DTMF on Twilio, browser, or raw WebSocket calls | Not supported | +| SIP INFO DTMF | Not supported | +| Incoming keypad collection | Not supported. Leave `keypadInputPlan.enabled` disabled | + +### Tools + +| Tool type | Support | +| --- | --- | +| `function`, `apiRequest`, `mcp` | Supported. Called by the reasoner | +| `endCall` | Supported | +| `transferCall`, `dtmf` | Supported on the connections listed above | +| Other tool types, including the Query tool, handoff, SMS, voicemail, and integration tools | Not supported. Use a function, API request, or MCP tool instead | -In **Reasoner**, enter: +Give tools unique names. Classic tool messages, such as request-start messages, aren't spoken. Guide what the assistant says while work runs in the speaker prompt. -```text title="Reasoner instructions" -Use lookupAvailability after the caller provides an explicit date -and time zone. Ask for clarification if either is missing or ambiguous. -Return only slots from the current lookup. Explain failures without -inventing results. No booking tool is available. -Call endCall when the caller asks to finish. -``` - - -Create a Function tool named `lookupAvailability` with the description “Check available appointment slots. Does not reserve or book a slot.” Add these required arguments: +### Assistant features -| Argument | Type | Description | -| --- | --- | --- | -| `date` | String | Requested calendar date in YYYY-MM-DD format | -| `timeZone` | String | IANA time zone, for example `America/New_York` | - -Set the tool's Server URL to your HTTPS handler. - -Attach the Function tool and an `endCall` tool to the assistant. You can also create a saved tool through [Create Tool](/api-reference/tools/create) and select it in the dashboard. - - -Set the first message to “Hello. I can check appointment availability. What date and time zone should I check?” and configure the assistant to speak first. - -Select **Publish**, review the changes, and complete the publication flow. Keep the assistant ID for API calls. - - - - - - -Save the following as `availability-assistant.json`: - -```json title="availability-assistant.json" -{ - "name": "Appointment availability assistant", - "model": { - "provider": "openai", - "model": "gpt-live-1", - "speaker": { - "instructions": "You help callers check appointment availability. Speak briefly. Collect an explicit calendar date and time zone, then delegate availability checks to the reasoner. If the caller corrects the date, delegate the corrected request and do not present old results as current. Report only confirmed lookup results. Availability does not reserve a slot. Booking is not connected. Delegate requests to end the call." - }, - "reasoner": { - "model": "gpt-5.6-terra", - "reasoningEffort": "low", - "instructions": "Use lookupAvailability only after the caller supplies an explicit calendar date and time zone. Ask for clarification if either is missing or ambiguous. Report only slots returned for the current request. Ignore superseded lookup results. A successful lookup does not reserve or book anything. Explain lookup failures without inventing slots. No booking tool is available. Call endCall when the caller asks to finish." - }, - "tools": [ - { - "type": "function", - "function": { - "name": "lookupAvailability", - "description": "Check available appointment slots. Does not reserve or book a slot.", - "parameters": { - "type": "object", - "properties": { - "date": { - "type": "string", - "description": "Requested calendar date in YYYY-MM-DD format." - }, - "timeZone": { - "type": "string", - "description": "IANA time zone, such as America/New_York." - } - }, - "required": [ - "date", - "timeZone" - ] - } - }, - "server": { - "url": "https://api.example.com/vapi/tools" - } - }, - { - "type": "endCall" - } - ] - }, - "voice": { - "provider": "openai", - "voiceId": "marin" - }, - "firstMessageMode": "assistant-speaks-first", - "firstMessage": "Hello. I can check appointment availability. What date and time zone should I check?", - "maxDurationSeconds": 600 -} -``` - - - -Set `VAPI_PRIVATE_API_KEY` to your Vapi private API key, then create the assistant: +| Feature | Support | +| --- | --- | +| Squads and assistant handoffs | Not supported. See [Migrate to GPT-Live](/gpt-live/migrate#from-a-squad) | +| Knowledge bases | Not available to the reasoner. Use a retrieval tool | +| Exact or prerecorded speech | Not supported. All speech, including greetings and idle check-ins, is generated | +| Audio URL greetings, generated first-message mode | Not supported | +| Idle-message hooks | Supported with `say.prompt`, `message.add`, and inline `endCall`. See [Idle messages](#idle-messages) | +| Other assistant hook events | Not covered by idle-message support. Check before relying on them | +| Live call control | End call, append speaker context, and cold transfer over HTTP. See [Live call control](#live-call-control) | +| Voice Simulations | Supported in Voice mode with a single assistant. See [Test and improve](/gpt-live/testing#repeat-scenarios-with-voice-simulations) | -```bash -curl --fail-with-body https://api.vapi.ai/assistant \ - -H "Authorization: Bearer $VAPI_PRIVATE_API_KEY" \ - -H "Content-Type: application/json" \ - --data-binary @availability-assistant.json -``` +### Recordings and post-call data -Keep the returned `id`. To change an existing assistant, use [Update Assistant](/api-reference/assistants/update), preserving its other tools and settings. - - - - +GPT-Live supports post-call analysis, structured outputs, scorecards, Boards, and Monitoring. Availability depends on your configuration and retained call data: -### Make a test call +| Setting or feature | Behavior | +| --- | --- | +| Recording disabled in the assistant | No recording is created | +| Recording-consent plan configured | Recording is disabled, because GPT-Live doesn't collect that consent | +| Transcript-based analysis | Requires transcripts | +| Audio-based analysis | Requires a recording | +| Monitoring with Zero Data Retention | Not supported | +| Live listen or monitor sockets | Not supported. Post-call Monitoring doesn't provide live audio | -Open the saved assistant in the dashboard and select **Talk**. Allow microphone access, then ask for availability on an explicit date in a specific time zone. +### Settings that no longer apply -Check that your handler receives the right values and the assistant reads back the returned slots. Ask to book a time. It should explain that booking isn't connected. Finish with “Please end the call” and confirm that the call ends. +Classic transcriber selection, text-to-speech settings, and start and stop speaking plans don't configure a GPT-Live conversation. General model fields such as `temperature` and `maxTokens` don't tune it either. ## Model settings @@ -164,43 +89,47 @@ These fields belong to the assistant object. See [Create Assistant](/api-referen | --- | --- | | `model.provider` | `openai` | | `model.model` | `gpt-live-1` | -| `model.speaker.instructions` | Speaker prompt. See [prompt behavior](/gpt-live/configuration#prompt-defaults-and-overrides). | -| `model.speaker.personalityPacks` | Array of `eager-listener`, `idle-hummer`, `bouncy`, and `unhurried`. | -| `model.reasoner.provider` | `openai`. Defaults to OpenAI when omitted. | -| `model.reasoner.model` | `gpt-5.6-sol`, `gpt-5.6-terra`, or `gpt-5.6-luna`. Defaults to `gpt-5.6-terra`. | -| `model.reasoner.reasoningEffort` | `none`, `low`, `medium`, `high`, `xhigh`, or `max`. Defaults to `low`. | -| `model.reasoner.instructions` | Custom reasoner prompt. Omission uses Vapi's default. | -| `model.tools` | Inline supported tool definitions. | -| `model.toolIds` | IDs of saved tools to resolve for the assistant. | +| `model.speaker.instructions` | Speaker prompt. See [prompt defaults](#prompt-defaults-and-overrides) | +| `model.speaker.personalityPacks` | Array of `eager-listener`, `idle-hummer`, `bouncy`, and `unhurried` | +| `model.reasoner.provider` | `openai`. Defaults to OpenAI when omitted | +| `model.reasoner.model` | `gpt-5.6-sol`, `gpt-5.6-terra`, or `gpt-5.6-luna`. Defaults to `gpt-5.6-terra` | +| `model.reasoner.reasoningEffort` | `none`, `low`, `medium`, `high`, `xhigh`, or `max`. Defaults to `low`. Higher effort can increase response time | +| `model.reasoner.instructions` | Reasoner prompt. Omit to use Vapi's default | +| `model.tools` | Inline tool definitions available to the reasoner | +| `model.toolIds` | IDs of saved tools | | `voice.provider` | `openai` | -| `voice.voiceId` | One of the [22 supported voices](/gpt-live/configuration#voices). | +| `voice.voiceId` | One of the [22 supported voices](#voices) | ## Prompt defaults and overrides -Set `model.speaker.instructions` for the spoken conversation and `model.reasoner.instructions` for task procedures. See [where instructions belong](/gpt-live/design#where-should-an-instruction-go), the [paired prompt example](/gpt-live/design#example-split-a-scheduling-prompt), and [what context the reasoner receives](/gpt-live/design#what-context-does-the-reasoner-receive). - -- Explicit speaker instructions take precedence over `model.systemPrompt` and system messages in `model.messages`. If you omit the speaker instructions, Vapi uses that legacy configuration. -- Omitting reasoner instructions uses Vapi's default reasoner prompt. Custom instructions replace that prompt entirely. +- Explicit speaker instructions take precedence over `model.systemPrompt` and system messages in `model.messages`. If you omit the speaker instructions, Vapi uses that classic configuration. +- Omitting reasoner instructions uses Vapi's default reasoner prompt. Custom instructions replace that prompt entirely. Vapi doesn't append behavioral instructions to a custom prompt. - An explicit empty string stays empty. Updating one prompt doesn't update the other. - -Personality packs append guidance to the speaker prompt. Set `model.speaker.personalityPacks` to an array of the IDs above and test one pack at a time. +- Personality packs append guidance to the speaker prompt. ## Personality and language -Personality packs append speaking-style guidance to your speaker prompt. Add this inside the assistant's `model` object: +Set `model.speaker.personalityPacks` to an array of pack IDs. Each pack appends speaking-style guidance to the speaker prompt: + +| Pack | What its guidance encourages | +| --- | --- | +| `eager-listener` | Brief acknowledgments at natural openings, with room for the caller to continue | +| `idle-hummer` | Soft humming or wordless sounds during quiet moments while waiting for tools | +| `bouncy` | More expressive pitch, emphasis, and rhythm, adapting to the caller's mood | +| `unhurried` | A slower pace, clear articulation, and space between ideas | ```json { "speaker": { - "instructions": "Speak warmly and briefly. Slow down for appointment dates and times. Follow requests to speed up or explain more. Delegate requests that need tools.", + "instructions": "Speak warmly and briefly. Slow down for dates and times.", "personalityPacks": ["unhurried"] } } ``` -Available IDs are `eager-listener`, `idle-hummer`, `bouncy`, and `unhurried`. See [Speaking style](/gpt-live/design#speaking-style) for what each pack encourages and how to combine personality with pacing. These are prompt instructions, not fixed speed controls. +Packs are prompt guidance, not fixed speed or pause controls. Test one at a time. See [Speaking style](/gpt-live/design#speaking-style) for how to choose. -Write the speaker prompt and `firstMessage` in the language you want the assistant to use. For Bossa or Tempo, you can include: +Write the speaker prompt and `firstMessage` in the language you want the assistant to speak. For Bossa or Tempo, for example: ```text Fale em português do Brasil, com respostas curtas e claras. @@ -208,7 +137,7 @@ Faça uma pergunta por vez. Só confirme um agendamento quando a ferramenta confirmar que ele foi realizado. ``` -Keep tool and delegation instructions in the prompt too. Try names, dates, and product terms in real calls, and add pronunciation guidance where needed. +Keep delegation triggers in the prompt in the same language. Try names, dates, and product terms on real calls, and add pronunciation guidance where needed. ## Voices @@ -235,8 +164,6 @@ Use the tags to narrow your choices, then listen to the samples. Bossa and Tempo ### More voices -You can also choose any of these voices. Select a preview to listen. - | Voice | API ID | Preview | | --- | --- | --- | | Alloy | `alloy` | [Listen to Alloy](https://voices-storage.vapi.ai/gpt-live/marble-alpha-v6/alloy.mp3) | @@ -252,162 +179,67 @@ You can also choose any of these voices. Select a preview to listen. ## Greetings and duration -Use a text greeting if the assistant should begin the conversation. - | Field or feature | Behavior | | --- | --- | -| `firstMessageMode: "assistant-speaks-first"` | Speaks the configured text greeting. | -| `firstMessageMode: "assistant-waits-for-user"` | Waits for the caller. | -| `firstMessage` | Text greeting. Without usable text, the assistant waits. | -| `maxDurationSeconds` | Maximum call duration in seconds. | -| Audio URL greeting | Unsupported. Not played as a greeting. | -| Generated-message first-message mode | Unsupported. The assistant waits for the caller. | +| `firstMessageMode: "assistant-speaks-first"` | Speaks a greeting generated from the text `firstMessage` | +| `firstMessageMode: "assistant-waits-for-user"` | Waits for the caller | +| `firstMessage` | Text greeting. Without usable text, the assistant waits | +| `maxDurationSeconds` | Maximum call duration in seconds | +| Audio URL greeting | Not supported. Not played as a greeting | +| Generated first-message mode | Not supported. The assistant waits for the caller | -Generated speech may vary from the supplied greeting. A text greeting is not a guarantee of exact prerecorded playback. +The spoken greeting may differ from the text you supply. ## Idle messages -Configure idle messages with `customer.speech.timeout` [assistant hooks](/assistants/assistant-hooks). GPT-Live can check in when the caller hasn't responded, make another check-in later, and optionally end the call with a final hook. - -Use `say.prompt` to guide each idle message. GPT-Live generates the spoken response; it doesn't guarantee exact wording or verbatim playback. - -Add these fields at the top level of your assistant configuration, preserving any existing hooks: - -```json title="Two idle check-ins and an optional final hangup" -{ - "hooks": [ - { - "on": "customer.speech.timeout", - "options": { - "timeoutSeconds": 8, - "triggerMaxCount": 3, - "triggerResetMode": "onUserSpeech" - }, - "do": [ - { "type": "say", "prompt": "Briefly ask whether the caller is still there." } - ] - }, - { - "on": "customer.speech.timeout", - "options": { - "timeoutSeconds": 16, - "triggerMaxCount": 3, - "triggerResetMode": "onUserSpeech" - }, - "do": [ - { "type": "say", "prompt": "Briefly ask whether the caller would like more time." } - ] - }, - { - "on": "customer.speech.timeout", - "options": { - "timeoutSeconds": 30, - "triggerMaxCount": 1 - }, - "do": [ - { "type": "say", "prompt": "Say a brief goodbye because the caller hasn't responded." }, - { "type": "tool", "tool": { "type": "endCall" } } - ] - } - ] -} -``` - -The first two hooks request check-ins after 8 and 16 seconds. At 30 seconds of continued caller silence, the optional third hook requests a farewell, then ends the call. All three delays use the same silence window: a check-in doesn't postpone the later hooks. Each hook fires at most once per window. To add another check-in within that window, add a separate hook with a different timeout. +Idle messages use `customer.speech.timeout` [assistant hooks](/assistants/assistant-hooks). See [When the caller goes quiet](/gpt-live/design#when-the-caller-goes-quiet) for a complete example. | Hook option | Behavior | | --- | --- | -| `options.timeoutSeconds` | Set explicitly for each hook. Accepts 2–1000 seconds. | -| `options.triggerMaxCount` | Limits triggers across silence windows. Accepts 1–10; defaults to 3. | -| `options.triggerResetMode` | Defaults to `never`, keeping the trigger count across the call. Set `onUserSpeech` to reset the count when the caller speaks. | +| `options.timeoutSeconds` | Set explicitly for each hook. Accepts 2–1000 seconds | +| `options.triggerMaxCount` | Limits how many times the hook fires across quiet periods. Accepts 1–10. Defaults to 3 | +| `options.triggerResetMode` | Defaults to `never`, which keeps the count for the whole call. Set `onUserSpeech` to reset the count when the caller speaks | + +How timing works: + +- A quiet period starts when the conversation becomes active, and again after the assistant speaks or the reasoner finishes work. +- Activity is measured from speech that appears in the transcript. Background noise doesn't count. +- All hooks count from the start of the same quiet period. A check-in doesn't postpone later hooks, and each hook fires at most once per quiet period. For another check-in in the same quiet period, add a separate hook with a different `timeoutSeconds`. +- When the caller speaks, check-ins stop until the next quiet period starts. A check-in that was already being spoken may still be heard. +- If the caller asks for a moment and the assistant replies, that reply starts a new quiet period. +- Check-ins wait while the reasoner is actively working. Once an async tool has been dispatched, check-ins can resume even though its result is still pending. -The first silence window starts when the conversation becomes active. Caller speech cancels the remaining check-ins for that window; an ordinary assistant response or further reasoner activity can start a new window. The hook's own speech doesn't rearm it. Pending reasoner work and transfers defer silence handling. Timing follows conversation activity, so test the delays with your prompts and tools. +Timing follows conversation activity, so test the delays with your prompts and tools. ### Supported idle-message hook actions | Action | GPT-Live behavior | | --- | --- | -| `say` with `prompt` | Generates one spoken response guided by the prompt. The model chooses the wording. | -| `message.add` | Adds context to the speaker and requests a response by default. Set `triggerResponseEnabled: false` to add context without requesting speech. System and developer messages become speaker instructions; other roles become speaker context. | -| `tool` with inline `tool: { "type": "endCall" }` | Ends the call, allowing an accompanying spoken action to play first. Without a spoken action, ends immediately. | - -Other actions, including saved `toolId` references, function calls, and transfers, aren't executed by these GPT-Live hooks. Hook messages affect the speaker, not the reasoner's history. These hook actions are separate from HTTP live-call controls: HTTP `say` and `add-message` remain unsupported. - -Use `say.prompt` for a one-time check-in. Its speech request applies to one response; if the caller interrupts, the speaker is instructed to answer the caller without repeating or resuming the check-in. - - -An inline `endCall` action is an explicit instruction to hang up. Once that hook fires, caller speech doesn't cancel the pending hangup. Choose the timeout accordingly. - - -Test a silent caller, a caller returning after multiple check-ins, and a slow tool response. Confirm that check-ins stop when the caller returns and don't repeat during the resumed conversation. If you include an `endCall` hook, verify its final hangup behavior too. - -## Connect your tools +| `say` with `prompt` | Generates one spoken response guided by the prompt. The model chooses the wording. If the caller interrupts, it answers the caller without resuming the check-in | +| `message.add` | Adds context to the speaker and requests a response by default. Set `triggerResponseEnabled: false` to add context without requesting speech. System and developer messages become speaker instructions. Other roles become speaker context | +| `tool` with inline `tool: { "type": "endCall" }` | Ends the call, after an accompanying `say` action if there is one. Once the hook fires, caller speech doesn't cancel the hangup | -Handle Vapi's [`tool-calls` server message](/server-url/events) at your HTTPS endpoint. Read `message.toolCallList`, validate each function's arguments, and query your availability service. - -Return a result for each request using its actual `toolCallId`: - -```json -{ - "results": [ - { - "name": "lookupAvailability", - "toolCallId": "TOOL_CALL_ID_FROM_REQUEST", - "result": "There are openings at 2 PM and 4 PM on the requested date in the requested time zone. No appointment has been booked." - } - ] -} -``` - -Replace the sample result with your service's data. If the lookup fails, return the failure so the assistant can explain it. See [Server authentication](/server-url/server-authentication) to protect the endpoint. - -Supported tool types are `function`, `apiRequest`, `mcp`, `endCall`, `dtmf`, and `transferCall`. Give tools unique resolved function names. Save reusable tools and attach them through `model.toolIds`, or include definitions in `model.tools`. - -Use function, API request, or MCP tools to connect external data sources. For phone actions, check [transfer and keypad support](/gpt-live/limitations#transfers-and-keypad-input). - -### MCP tools - -You can keep the MCP tools you already use with an existing Vapi assistant when migrating to GPT-Live. Attach the same saved tool IDs through `model.toolIds`, or keep your inline `type: "mcp"` definitions in `model.tools`. Existing server URLs, headers, credentials, and protocol settings are reused. See [MCP tools](/tools/mcp) for setup. - -For example, add your saved MCP tool ID to the assistant's `model` configuration, preserving any other tool IDs: - -```json -{ - "toolIds": ["YOUR_EXISTING_MCP_TOOL_ID"] -} -``` - -At call setup, Vapi connects to the MCP server and discovers its tools. Those tools are available to the **reasoner**. The speaker delegates work to the reasoner, which calls the tool and returns the result for the conversation. Put tool-use procedures in `model.reasoner.instructions` and delegation guidance in `model.speaker.instructions`. - -The discovered tool list stays fixed for that call. Each tool invocation opens a separate MCP connection; it doesn't rediscover the tool list. Start a new call to pick up changes to the server's available tools. GPT-Live doesn't play classic per-tool speech messages; use speaker instructions to guide how it talks while work runs. - -### Slow and asynchronous requests - -A synchronous function waits for your handler's result. Set `async: true` on a function tool to let the conversation continue while that request completes. Return the final result through the original request, using the matching `toolCallId`. - -If your handler returns “queued” before the work finishes, add a status lookup tool. Tell the reasoner when to check it. A later callback doesn't automatically update the conversation. - -Tools can run concurrently. Keep dependent actions in order: check availability, confirm the caller's choice, then book. Your service should validate those requirements too. +Other actions, including saved `toolId` references, function calls, and transfers, aren't run by GPT-Live hooks. The `endCall` tool's own messages aren't used, so configure the goodbye as a separate `say` action. Hook messages affect the speaker, not the reasoner. Hook actions are separate from HTTP live call control. ## Connect a call -Start with dashboard **Talk**, then connect your intended call channel. GPT-Live supports browser WebRTC, native Twilio, Vapi SIP, and raw WebSocket audio. See [connection limits](/gpt-live/limitations#calls-and-voice) before changing an existing setup. +Start with the dashboard's **Talk** button, then connect your call channel. GPT-Live supports browser WebRTC, native Twilio, Vapi SIP, and raw WebSocket audio. ### Phone calls For inbound Twilio calls, [import a Twilio phone number](/phone-numbers/import-twilio) and assign your assistant to it. For SIP, follow the [SIP guide](/advanced/sip). -For an outbound Twilio test, send this body to [Create Call](/api-reference/calls/create): +For an outbound Twilio test, send this body to [Create Call](/api-reference/calls/create), using a customer number you control: ```json { "assistantId": "YOUR_ASSISTANT_ID", "phoneNumberId": "YOUR_TWILIO_PHONE_NUMBER_ID", - "customer": {"number": "+14155550100"} + "customer": { "number": "+14155550100" } } ``` -Replace the customer phone number with one you control. `phoneNumberId` is the Vapi ID of the imported Twilio phone number. +`phoneNumberId` is the ID Vapi gave your imported Twilio number, not Twilio's own phone number SID. ### WebSocket configuration @@ -427,15 +259,13 @@ Send this body to `POST https://api.vapi.ai/call` with your Vapi private API key } ``` -Connect to `transport.websocketCallUrl` in the response. Send and receive binary audio frames containing signed 16-bit little-endian samples at 24 kHz. Convert your microphone audio to that format before sending it. +Connect to `transport.websocketCallUrl` in the response. Send and receive binary frames of mono, signed 16-bit little-endian audio at 24 kHz. The default format is rejected, so set `sampleRate` to `24000` explicitly. Browser and phone connections negotiate their own formats. -Set `sampleRate` to `24000` for GPT-Live. Browser WebRTC and phone connections negotiate their own audio formats. - -Close the WebSocket to end the call, have the assistant invoke `endCall`, or send an HTTP `end-call` request to the [live control URL](#live-call-control). Send control commands over HTTP, not as JSON frames on the audio WebSocket. See [WebSocket transport](/calls/websocket-transport) for the general transport, applying the restrictions above. +To end the call, close the WebSocket, have the assistant call `endCall`, or send an HTTP `end-call` request to the [live control URL](#live-call-control). Send controls over HTTP, not as JSON frames on the audio WebSocket. See [WebSocket transport](/calls/websocket-transport) for the general transport. ## Live call control -Your application can end an active GPT-Live call, append speaker context, or initiate a cold transfer. Send an HTTP POST to that call's `monitor.controlUrl`, returned in the call object. These commands aren't supported through audio WebSocket text frames or SDK data channels. +Your application can end an active GPT-Live call, add context for the speaker, or start a cold transfer. Send an HTTP POST to the call's `monitor.controlUrl`, returned in the call object. These commands aren't accepted as audio WebSocket text frames or SDK data-channel messages. Set `CONTROL_URL` to the returned URL. Call control must be enabled in the assistant's `monitorPlan`: setting `controlEnabled` to `false` disables it. If `controlAuthenticationEnabled` is `true`, include `Authorization: Bearer YOUR_VAPI_PUBLIC_API_KEY` from the same organization on each request. The examples below assume control authentication is disabled. Keep the control URL private. @@ -453,9 +283,9 @@ Use `append-context` with a `kind` and nonblank `content` of up to 16,000 charac | Kind | Purpose | Example content | | --- | --- | --- | -| `commentary` | Give the speaker information to convey in its own words | `Your appointment is confirmed for Tuesday at 2 PM.` | -| `thinking` | Add silent context for the speaker to use | `The caller has already verified their account.` | -| `instructions` | Guide the speaker's behavior or ask it to try saying something | `Tell the caller their order is ready, then ask whether they need anything else.` | +| `commentary` | Information for the speaker to convey in its own words | `Your appointment is confirmed for Tuesday at 2 PM.` | +| `thinking` | Context for the speaker to use, without prompting speech | `The caller has already verified their account.` | +| `instructions` | Guidance for the speaker's behavior, including asking it to try saying something | `Tell the caller their order is ready, then ask whether they need anything else.` | ```bash curl --fail-with-body -X POST "$CONTROL_URL" \ @@ -463,13 +293,13 @@ curl --fail-with-body -X POST "$CONTROL_URL" \ -d '{"type":"append-context","kind":"instructions","content":"Tell the caller their order is ready, then ask whether they need anything else."}' ``` -These messages go to the speaker. They don't add context to the reasoner's history or directly start reasoning or tool work. `thinking` doesn't request speech, but it isn't a separate private context store: the speaker can use that information in subsequent responses. +This context goes to the speaker only. It doesn't reach the reasoner or start reasoning or tool work, and it isn't added to the transcript as speech. `thinking` doesn't request speech, but the speaker can use it and repeat it later, so don't send secrets that way. -A successful append response means the context was submitted. It doesn't guarantee exact wording, immediate speech, or completed playback. +A successful response means the context was submitted. It doesn't guarantee exact wording, immediate speech, or completed playback. ### Cold transfer -On native Twilio and Vapi SIP calls, send a phone or SIP destination. For example: +On native Twilio and Vapi SIP calls, send a phone or SIP destination: ```bash curl --fail-with-body -X POST "$CONTROL_URL" \ @@ -477,42 +307,48 @@ curl --fail-with-body -X POST "$CONTROL_URL" \ -d '{"type":"transfer","destination":{"type":"number","number":"+14155550100","transferPlan":{"mode":"blind-transfer"}}}' ``` -Replace the number with your transfer destination. For SIP, use a destination such as `{"type":"sip","sipUri":"sip:agent@example.com"}`. An application supplies the destination in this request; a reasoner using a `transferCall` tool selects from that tool's configured destinations. +For SIP, use a destination such as `{"type":"sip","sipUri":"sip:agent@example.com"}`. Your application chooses the destination in this request. A reasoner using a `transferCall` tool chooses only from that tool's configured destinations. -Cold transfer isn't supported on browser or raw WebSocket calls. Don't include pre-transfer `content`, warm-transfer plans, generated summaries, or fallback plans. See [transfer restrictions](/gpt-live/limitations#transfers-and-keypad-input). Carrier acceptance doesn't establish that the destination answered. +Cold transfer isn't supported on browser or raw WebSocket calls. Don't include pre-transfer `content`, warm-transfer plans, generated summaries, or fallback plans. A successful response means the carrier accepted the transfer, not that the destination answered. ### Responses and unsupported controls -Successful commands return HTTP `200` with `{"status":"ok"}`. Rejected requests return an error, including when the call isn't active or is already transferring. If a `502` response reports an unconfirmed outcome, don't automatically retry: the action may have reached the provider. A `503` response stating that no action was taken is safe to retry. +Successful commands return HTTP `200` with `{"status":"ok"}`. Rejected requests return an error, including when the call isn't active or is already transferring. If a `502` response reports an unconfirmed outcome, don't retry automatically: the action may have reached the provider. A `503` response means no action was taken and is safe to retry. -GPT-Live doesn't support HTTP `say`, `add-message`, or mute/unmute controls. Use `append-context` for speaker guidance. The `say` and `message.add` actions in [idle-message hooks](#idle-messages) are configured on the assistant and aren't HTTP control commands. See [live control compatibility](/gpt-live/limitations#can-i-inject-commentary-or-control-speech-during-a-call). +GPT-Live doesn't support the HTTP `say`, `add-message`, or mute and unmute controls. Use `append-context` for speaker guidance. The `say` and `message.add` actions in [idle-message hooks](#idle-messages) are assistant configuration, not HTTP controls. -## Review and monitor calls +## Moved sections -GPT-Live supports post-call analysis, structured outputs, scorecards, Boards, and Monitoring. Configure the features you need, then inspect the results after a test call finishes. +These sections used to be on this page. -| Feature | Setup and results | -| --- | --- | -| [Structured outputs](/assistants/structured-outputs-quickstart) | Create output definitions, attach their IDs through `artifactPlan.structuredOutputIds`, and read `call.artifact.structuredOutputs`. Keep existing IDs and artifact settings when updating. | -| [Call analysis](/assistants/call-analysis) | Existing `analysisPlan` configurations work. Prefer structured outputs for new setups. | -| [Scorecards](/observability/scorecard-quickstart) | Grade boolean or numeric outputs and attach the scorecard to your assistant. Read `call.artifact.scorecards`. | -| [Boards](/observability/boards-quickstart) | Track available call metrics and outcome trends for your assistant. Classic transcriber and text-to-speech timing fields don't describe GPT-Live. | -| [Monitoring](/observability/monitoring-quickstart) | Configure supported conditions, thresholds, evaluation windows, and notification destinations. Requires retained data. Unavailable with Zero Data Retention. | -| [End-of-call reports](/server-url/events) | Configure a Server URL and the `end-of-call-report` server message. [Authenticate your endpoint](/server-url/server-authentication) and handle repeated delivery by call ID. | +### Before you start -Keep transcript collection enabled for transcript-based checks. Audio-based checks require a recording. See [recording and data restrictions](/gpt-live/limitations#recordings-and-post-call-data) before relying on those artifacts. +Prerequisites are now in the [Quickstart](/gpt-live/quickstart#before-you-start) and [Build with tools](/gpt-live/tools#before-you-start). -Retrieve a completed call with your private API key: +### Create an assistant -```bash -curl --fail-with-body "https://api.vapi.ai/call/$CALL_ID" \ - -H "Authorization: Bearer $VAPI_PRIVATE_API_KEY" -``` +See the [Quickstart](/gpt-live/quickstart#create-the-assistant) for a first assistant, and [Build with tools](/gpt-live/tools) for connecting tools. + +### Make a test call + +See [Talk to it](/gpt-live/quickstart#talk-to-it) and [Check what happened](/gpt-live/quickstart#check-what-happened). + +### Connect your tools + +See [Build with tools](/gpt-live/tools) and [the tool response contract](/gpt-live/tools#the-tool-response-contract). + +### MCP tools + +See [Use existing MCP tools](/gpt-live/tools#use-existing-mcp-tools). + +### Slow and asynchronous requests + +See [Slow and external work](/gpt-live/tools#slow-and-external-work). -Review the transcript and recording alongside tool results. A missing analysis result isn't a pass. Check attached output IDs, processing, and available artifacts. Creating a scorecard doesn't automatically configure a monitor. +### Review and monitor calls -When testing a migration, keep the previous assistant available so you can route new calls back. End test calls through the call interface, `endCall`, or by closing your WebSocket. Check any external work still running. Ending the call doesn't undo it. +See [Watch production calls](/gpt-live/testing#watch-production-calls). -## Cost +### Cost -Review Vapi call costs for your configuration. Reasoning, tools, and telephony can contribute to the total. OpenAI's [model pricing](https://developers.openai.com/api/docs/models/gpt-live-1) describes provider charges rather than the complete cost of a Vapi call. +See [Cost](/gpt-live/testing#cost). diff --git a/fern/gpt-live/design.mdx b/fern/gpt-live/design.mdx index 7ef28ad81..0c135c119 100644 --- a/fern/gpt-live/design.mdx +++ b/fern/gpt-live/design.mdx @@ -1,268 +1,585 @@ --- -title: Design your GPT-Live assistant -subtitle: Plan the conversation and choose how your assistant speaks -description: Design GPT-Live conversations with speaker and reasoner roles, outcome-based delegation, useful overlap during tools, flexible pacing, and personality packs. +title: Design GPT-Live conversations +subtitle: Help callers reach their goal while task work happens alongside the conversation +description: Design GPT-Live conversations around the caller's aim, with speaking and listening happening while the reasoner and tools work. Covers dependencies, useful waiting, returning results, corrections, speaking style, personality packs, idle callers, and paired prompts. slug: gpt-live/design --- -With GPT-Live, the conversation can continue while work happens in the background. A caller can add a preference during a lookup, ask a question halfway through scheduling, or ask the assistant to slow down. +A GPT-Live assistant keeps listening and speaking while its reasoner looks things up and takes actions. The conversation and the work run side by side, so part of the design is deciding how they fit together: which work to start early, what the conversation can usefully cover while it runs, and how a result rejoins the conversation when it arrives. -The design task is to make those moments feel connected. Decide what the assistant is trying to accomplish, how it should guide the caller, and which decisions need help from your services. +This page works through one conversation: a caller booking an appointment with a small service business. The same thinking applies to support, intake, and other calls that mix questions with actions. [Build with tools](/gpt-live/tools) covers the tools behind this example. -This guide uses scheduling as an example. The same approach applies to support, intake, and other conversations that combine questions with actions. +The dialogues on this page are illustrations of the behavior to design for. GPT-Live generates its own wording, and results vary between calls, so test your prompts on real calls before relying on a pattern. -**Jump to:** [Architecture](#speaker-and-reasoner-roles) · [Prompt example](#example-split-a-scheduling-prompt) · [Context](#what-context-does-the-reasoner-receive) · [Squads](#start-with-one-assistant) · [Conversation flow](#questions-and-changes-of-topic) · [Voice design](#speaking-style) · [Testing](#testing-conversations) +## Start with the caller's aim + +Before writing a prompt, decide what the caller is trying to get done and what a good ending sounds like. For an appointment, the call has succeeded when the caller has a confirmed time that suits them, the details are right, and they know what happens next. Filling in every field or sounding friendly isn't enough on its own. + +Then sort what the task needs: + +| Kind of information | For an appointment | +| --- | --- | +| Needed to look up availability | Service, location, date | +| Needed to book | A time from that lookup, the name for the booking, the caller's agreement | +| Helpful but optional | A time-of-day preference | +| Alternatives if nothing fits | Another day, the other location | + +This table does most of the design work. It tells you when work can start, which questions can wait, and what the assistant must not claim until a result arrives. + +Callers rarely supply details in the order you'd ask for them. An informed caller may give everything at once: + +> **Caller:** Hi, I'd like a consultation downtown this Friday, afternoon if possible. +> +> **Assistant:** Sure, I'll check Friday afternoon downtown. Who should I put the appointment under? + +The assistant already has the service, location, and date, so it starts the lookup and asks for the one detail that booking will need. Asking "Which service would you like?" here would make the caller repeat themselves. Tell the speaker to use details the caller has already given, and to ask for only the next missing piece. ## Speaker and reasoner roles -The **speaker** carries the spoken interaction: listening, responding, asking questions, and deciding when to delegate. The **reasoner** takes a delegated task, follows its instructions, and calls tools. A tool connects the assistant to information or actions in your service. +A GPT-Live assistant has two parts. The **speaker** is the voice the caller hears. It listens, speaks, and decides when to ask for help. The **reasoner** does the task work. It follows your procedures, calls your tools, and returns what it found. Asking the reasoner for help is called **delegation**. -Both models belong to the same assistant. The caller stays in one conversation while the reasoner works behind it. +The two run at the same time. Here's the opening of the appointment call, step by step: -| Give the speaker | Give the reasoner | -| --- | --- | -| The assistant's role and the caller's goal | The procedure for completing each task | -| Tone, language, and pacing | Required information and confirmation rules | -| When to clarify or delegate | Which tools to use and in what order | -| How to explain results to the caller | How to interpret tool results and failures | + + +**Conversation:** "A consultation downtown this Friday, afternoon if possible." -For example, the speaker needs to know that it can help find an appointment. The reasoner needs to know which service returns availability, what information that service requires, and when booking is allowed. +**Background work:** None yet. The speaker now has the service, location, and date, so it delegates the lookup. + + +**Conversation:** The assistant asks who the appointment is for, and the caller says "Sam Lee." -This separation lets you keep the speaker's instructions focused. Put detailed business procedures in the reasoner prompt. Your services remain responsible for enforcing permissions and validating actions. +**Background work:** The reasoner checks Friday's open times downtown. + + +**Conversation:** "Friday afternoon I have 2:30 or 4." -### Where should an instruction go? +**Background work:** The lookup returned 10:30, 2:30, and 4. The speaker offers the two that fit the afternoon request. + + -Use `model.speaker.instructions` for **how to conduct the conversation and when to ask for help**. Use `model.reasoner.instructions` for **how to decide and act once asked**. +Think of it as two lanes: the conversation the caller hears, and the work the reasoner and your tools are doing. Good design keeps both moving toward the caller's aim, and brings them together when a result matters. -| Instruction | Where it belongs | Why | +This concurrency is a property of the conversation, and it applies to every tool. It's separate from the `async` setting on a function tool, which controls whether the reasoner waits for your webhook's response before continuing. See [Slow and external work](/gpt-live/tools#slow-and-external-work) for that. + +### What context does the reasoner receive? + +The speaker and reasoner see different things. You need this to predict how the assistant handles changes. + +| | Speaker | Reasoner | | --- | --- | --- | -| “Ask one question at a time and speak briefly.” | Speaker | Controls the interaction the caller hears. | -| “Delegate availability checks and booking requests.” | Speaker | Tells it when to involve the reasoner. | -| “Use `lookupAvailability` with an explicit date and time zone.” | Reasoner | Defines how to use a tool correctly. | -| “Call `bookAppointment` only after the caller confirms a specific slot.” | Reasoner | Defines the condition for taking an action. | -| “Never tell the caller a booking succeeded before the tool confirms it.” | Both, tailored to each role | The reasoner must verify success; the speaker must describe it accurately. | +| Live audio | Hears the caller and speaks continuously | Doesn't hear audio. Works from the conversation transcript | +| Conversation | Is part of it | Receives the transcript available when the request starts | +| Instructions | Speaker prompt, plus any personality packs | Reasoner prompt. The speaker prompt isn't copied over | +| Tools | None. It asks the reasoner | Your tools. It can use its earlier work and tool results from the call | -The prompts are separate. Instructions written only in the speaker prompt aren't automatically included in the reasoner prompt. Put rules that affect both conversation and actions in both prompts, using wording appropriate to each role. Keep detailed tool procedures in the reasoner prompt rather than copying the entire prompt into both fields. +Two consequences shape the rest of this page. -### Example: split a scheduling prompt +First, the reasoner starts from a snapshot. If the caller says "Actually, Thursday" while a Friday lookup is running, that lookup doesn't change. The speaker hears the correction, and the next request to the reasoner will include it. Design for a new request when the details change. -Suppose the assistant can look up availability and book appointments. The example below assumes you have attached tools named `lookupAvailability` and `bookAppointment`. Replace those names and requirements with your actual tools. The [configuration walkthrough](/gpt-live/configuration#create-an-assistant) provides a lookup-only example if you haven't connected booking yet. +Second, a result is information for the speaker to use in the current moment. The speaker doesn't need to recite it. It can take the open times, match them against what the caller said while the lookup ran, and offer the ones that fit. -```text title="model.speaker.instructions" -You help callers find and book appointments. Speak warmly and briefly. -Ask one question at a time. Use details the caller has already supplied. +## Order the work by what depends on what + +Start each piece of work as soon as its inputs are known, and use the conversation for things that don't depend on it. Work through the dependencies for the appointment: + +- A lookup needs the service, location, and date. Once those are known, delegate right away. +- The booking name doesn't affect the lookup, so collect it while the lookup runs. +- A time can be chosen only after the open times are known. +- A booking can be announced only after the booking tool confirms it. -Collect an explicit date and time zone, then delegate availability checks. -Read back the returned options. Ask the caller to confirm one specific -slot before delegating a booking request. +The useful test for any question during a lookup is whether the outstanding work still answers the caller's need after they reply. In this example, the lookup tool returns every open time for the requested day and location. So if the caller says "afternoon, ideally after three" while that lookup runs, the result still covers it, and the assistant can filter. If the caller says "Actually, can we do Riverside instead?", the running lookup no longer answers the request, and a new one is needed. -While a lookup runs, you may ask about preferences that don't change -its inputs. If the caller changes the date or time zone, acknowledge -the correction and delegate the updated request. Don't present an old -lookup result as an answer to the new request. +That's why the example lookup returns the whole day rather than taking a time-of-day argument. When you design your own tools, return enough to answer the likely follow-up questions without another round trip, as long as the result stays short. -Only say an appointment is booked after the reasoner returns a confirmed -booking. If information is missing or a tool fails, explain the issue -and ask for the next detail needed. +### When to delegate + +The speaker only starts work when it decides to delegate. A missed delegation can leave an assistant that says "Let me check that" and then nothing happens, or says goodbye without ending the call. Make the triggers concrete and list them in the speaker prompt: + +```text title="Delegation triggers in the speaker prompt" +Delegate to the backend when: +- You know the service, location, and date. Ask for availability right away. +- The caller has clearly said yes to booking after you read back the day, + time, location, and name. +- The caller changes the service, location, date, or chosen time. +- The caller asks about services, what to bring, locations, hours, or policies. +- The backend asked for a detail and the caller has now given it. +- The caller asks to end the call or says goodbye. Always delegate this. + Saying goodbye does not end the call. + +Do not delegate to the backend when: +- You can answer from what was already said or from a result that still + answers the current request. +- You need a brief clarification first. ``` -```text title="model.reasoner.instructions" -Help the speaker complete the caller's scheduling request using the -conversation transcript and the available tools. +Two of these triggers are easy to forget. When the reasoner returns a question ("Which location?"), the caller's answer needs a new delegation, or the task stalls. And ending the call is a tool action, so the speaker has to delegate it like any other task. -For availability, require an explicit date and time zone before calling -lookupAvailability. If either is missing or ambiguous, return a concise -clarification for the speaker to ask. Don't invent missing arguments. +## While work is running -For booking, require the caller's explicit confirmation of a specific -slot returned by lookupAvailability. Call bookAppointment only when -those conditions and the tool's required arguments are satisfied. +While a result is pending, the conversation doesn't have to stop, and what it does instead is a design choice. The aim is to make the time useful to the caller. Sometimes that means asking something, sometimes it means answering something, and sometimes it means a short acknowledgment and a pause. -Use the latest request in the supplied transcript. Before taking a new -action, check prior tool results so you don't repeat a completed booking. -Don't retry an uncertain booking outcome without checking its status. +### Talking while tools run -Return the relevant date, time zone, slots, or confirmed booking details -for the speaker to explain. If a tool fails, report the failure and the -next useful step. Don't invent availability or claim success without a -successful tool result. -``` +When the next step doesn't depend on the pending result, the conversation can move ahead. Compare a serial version with one that uses the wait: -The speaker asks the caller for clarification; the reasoner returns what needs clarification. The reasoner doesn't speak to the caller directly. Your service should also enforce booking requirements: prompt instructions don't replace validation in the tool handler. +> **Serial** +> +> **Caller:** A consultation downtown on Friday. +> +> **Assistant:** Let me check. *(silence while the lookup runs)* I have 10:30, 2:30, and 4. Which would you like? +> +> **Caller:** 2:30. +> +> **Assistant:** And what name should I put it under? -### What context does the reasoner receive? +> **Using the wait** +> +> **Caller:** A consultation downtown on Friday. +> +> **Assistant:** I'll check Friday downtown. Who should I put the appointment under? +> +> **Caller:** Sam Lee. +> +> **Assistant:** Thanks, Sam. Friday I have 10:30, 2:30, or 4. -**At each delegation, the reasoner receives the available spoken conversation transcript so far**, including caller and assistant turns. It also receives its own instructions, available tool definitions, and retained history from prior reasoner work and tool results within that call. It doesn't rely on a short summary or a few selected lines supplied by the speaker. +The name question works during the lookup because the booking will need it and the answer can't change the lookup. When the caller picks a time, the assistant already has what it needs to confirm and book. -The speaker and reasoner still have different context: +The name reaches the booking because the reasoner's next request starts from the updated transcript, which now includes it. You don't need to pass it along separately. -| Context | Speaker | Reasoner | -| --- | --- | --- | -| Live audio | Listens and speaks continuously | Receives text, not the audio stream | -| Spoken conversation | Part of the live conversation | Receives the transcript available when delegation starts | -| Instructions | Speaker prompt and personality guidance | Reasoner prompt; the speaker prompt isn't automatically copied | -| Tool work | Receives results returned by the reasoner | Has tool definitions and retained reasoner/tool history | -| HTTP `append-context` | Receives the injected speaker context | Doesn't receive it directly in its history | +Guide this in the speaker prompt: -The transcript is a snapshot of what Vapi has received at that point. Speech that arrives while the reasoner is already working doesn't automatically refresh that delegation's transcript. A subsequent delegation receives the updated transcript. Completed work and tool results can inform later reasoning, but the reasoner isn't continuously listening alongside the speaker. +```text +While work is running: +- If there is a useful question that doesn't change the request, ask it. + Collect the name for the booking while availability is checked. +``` -For example, if the caller says “Actually, Thursday” while a Friday lookup is running, the speaker can hear and acknowledge the correction. The ongoing lookup may still return Friday's availability. Delegate the changed request, match results to the date they answer, and don't assume an interruption cancels an action already sent to a tool. +Choose the question for what it contributes. A question the task doesn't need, such as asking for an email address the booking never uses, fills the wait but makes the call longer and asks the caller for more personal information than necessary. -### When to delegate +### When the next step is blocked -Give the speaker a clear reason to ask for help: “Delegate requests to find an appointment that fits the caller's preferences.” Give the reasoner the steps needed to produce that result. +Sometimes nothing useful can move forward until the result arrives. The caller has confirmed a time and the booking is running. The conversation can still be helpful, depending on the moment: -| Layer | Example instruction | -| --- | --- | -| Speaker | “Help the caller find a suitable appointment. Delegate availability checks and booking requests.” | -| Reasoner | “Confirm the date and time zone. Look up available slots. Book only after the caller confirms a specific slot.” | -| Service | Validate the requested slot and create the booking only when the request meets your requirements. | +- **Explain what's already known.** Tell the caller about preparation or a policy that applies whatever the result, such as arriving a few minutes early. +- **Answer a related question** from information the assistant already has. +- **Say what the result will let the caller decide**, if that helps them get ready. For example, "Once that's booked I'll give you a booking number to keep." +- **Invite a relevant question**, when the caller is likely to have one. +- **Acknowledge the wait once and pause**, when nothing useful remains or the caller would clearly prefer quiet. -Keep the conditions for delegation concrete. A greeting or a request to repeat an already confirmed time may stay with the speaker. A fresh lookup, changed booking request, or end-call request needs the reasoner and its tools. +Here the caller asks a question during the booking: -The reasoner's result gives the speaker information for its next response. The speaker can explain that information in the context of what the caller just said. See [Configure GPT-Live](/gpt-live/configuration#create-an-assistant) for an assistant with both prompts and a lookup tool. +> **Assistant:** That's Friday at 2:30, downtown, under Sam Lee. Shall I book it? +> +> **Caller:** Yes, please. +> +> **Assistant:** Booking that now. +> +> **Caller:** Do I need to bring anything? +> +> **Assistant:** For a consultation, nothing's required. Bring any questions you'd like to cover, and it helps to arrive five minutes early. +> +> **Assistant:** You're all set: Friday at 2:30 downtown, under Sam Lee. Your booking number is TEST-4F2A1C. -## Start with one assistant +The answer about what to bring doesn't depend on the booking, so the assistant can give it while the booking runs. It waits for the result before saying "You're all set." -Early testing suggests GPT-Live handles many of the needs that previously led builders to squads out of the box. The speaker maintains the conversation while the reasoner handles detailed instructions and task sequencing. You may be able to cover the full use case with one assistant. +The speaker can only answer from what it knows. To make answers like this possible during a wait, put a few stable facts the caller often asks about in the speaker prompt, as the [example prompt](#example-split-a-scheduling-prompt) does with arrival and preparation guidance. Anything that changes, or that you'd want looked up, belongs behind a tool. -Suppose your scheduling squad has separate assistants for intake, availability, and booking. Start by carrying those responsibilities into one GPT-Live assistant: +When nothing useful remains, a short acknowledgment and a pause is the right choice: -```mermaid -flowchart TD - A[One scheduling conversation] - A --> B[Speaker guides the caller] - B --> C[Reasoner manages the task] - C --> D[Collect required details] - D --> E[Look up availability] - E --> F[Book after confirmation] +> **Assistant:** Friday at 2:30, downtown, under Sam Lee. Shall I book it? +> +> **Caller:** Yes, the 2:30. +> +> **Assistant:** Booking that now. +> +> *(a short pause)* +> +> **Assistant:** You're booked for Friday at 2:30. + +Avoid the patterns that make waiting worse: + +- repeated "still checking" messages +- invented progress or time estimates ("almost done", "about ten more seconds") +- questions that don't serve the task, especially ones that collect more personal information +- talking over a caller who is thinking, or who asked for a moment +- any wording that implies the action succeeded before the result arrives + +```text +While work is running: +- If the caller asks something you can answer from information already given, + answer it. +- Otherwise acknowledge the wait once and give the caller room. Don't repeat + that you are still checking and don't guess how long it will take. ``` -This diagram shows responsibilities, not a fixed script for every call. The caller may supply several details at once or ask a question between stages. +## Bring the result back into the conversation -Preserve the outcomes and business rules you need. Test whether the new architecture still needs the divisions you created for the old one. Pay attention to missed steps and recovery from errors before consolidating a production flow. +Results arrive when the work finishes, which may not be a good moment in the conversation. The caller may be mid-sentence, or may have changed their mind while the work ran. The speaker's job is to fit the result into the conversation as it now stands. -Existing squad configurations and assistant handoffs are [currently unsupported](/gpt-live/limitations#can-i-use-an-existing-squad). The early finding is that some use cases need fewer assistants, not that every squad can be converted automatically. +### Confirming results -## Questions and changes of topic +Suppose the lookup finishes while the caller is adding a preference: -Give the conversation a destination and a few meaningful stages: understand the request, find suitable options, confirm the choice, and explain the result. The caller doesn't need to answer questions in the order you wrote them. +> **Assistant:** I'll check Friday downtown. Who should I put it under? +> +> **Caller:** Sam Lee. And actually, anything after three would be best. +> +> **Assistant:** After three on Friday, I have 4 o'clock. Would that work? -Consider this opening: +The result came back with 10:30, 2:30, and 4. The assistant let the caller finish, applied the new preference to the result it already had, and offered the one time that fits. It didn't read out all three, and it didn't ask "What time works for you?" after the caller had just said. -> **Caller:** I need something Friday afternoon. I've been there before. +Prompt for the behaviors you want when a result arrives: -The caller has already supplied a day, a preference, and a useful detail. Design the assistant to acknowledge those details and ask for the next missing piece, such as the exact date or time zone. Repeating questions the caller has answered makes even responsive speech feel mechanical. +- Let the caller finish before bringing up the result. +- Check the result against the caller's latest request. If it no longer answers it, delegate again. +- Present only what's relevant now. +- Don't ask a question the caller has just answered. +- State an outcome, such as a booking, only when the result confirms it. -A detour can help the caller make the next decision: +A prompt can't guarantee exact timing, so check these behaviors on real calls. When a result is out of date, the fix is a new delegation. The speaker shouldn't try to adjust a stale result by guessing. + +## Corrections, interruptions, and detours + +Callers change their minds, interrupt, and ask unrelated questions. These look similar in a transcript but need different handling. + +### When callers change the request -> **Caller:** Before we pick a time, how long does the appointment take? +Three things can sound alike: + +- **"Hold on, stop."** The caller wants the assistant to stop talking. Nothing about the task changes. +- **"Actually, Thursday."** The request has changed. New work is needed. +- **"Cancel my appointment."** The caller wants an action taken. The reasoner needs to call the cancel tool. + +Interrupting the assistant's speech doesn't stop work that's already running. A lookup or booking your service has received carries on. Here's a date correction during a lookup: + +> **Caller:** Friday downtown, please. +> +> **Assistant:** I'll check Friday downtown. > -> **Assistant:** Let me check that for you. Then we can find a time that works. +> **Caller:** Actually, sorry, Thursday. +> +> **Assistant:** No problem, Thursday instead. Checking that now. +> +> **Assistant:** On Thursday downtown I have 9 or 1. -If the duration is already available in trusted context, the speaker can explain it. If it needs a fresh lookup, delegate that question. Then return to scheduling with the caller's earlier preferences intact. +The Friday result may still come back. The speaker should recognize it no longer answers the request and use the Thursday one. Your tools can help by including the date and location in every result, as the example's tools do, so a stale result is easy to spot. -A useful speaker instruction is: +If the change arrives after an action has completed, the assistant should say so and follow your procedure for changing it: -```text title="Keep the conversation connected" -Use details the caller has already provided. Ask one question for the -next missing piece. If the caller asks a related question, address it -and then return to the unfinished task. Delegate questions that need -new information from our services. -``` +> **Caller:** Oh wait, can we make that Thursday instead? +> +> **Assistant:** Friday at 2:30 is already booked. I can look at Thursday, and if there's a time you like, move it. Shall I check? -## Talking while tools run +The booking exists until your service cancels or changes it. Tell the reasoner how to change a completed action, and have your service enforce it. In this example, a move is two separate steps: cancel the Friday booking, then book Thursday. So the procedure finds a Thursday time and gets the caller's agreement first, then cancels and books. If the second step fails, the assistant says what actually happened: Friday was cancelled and Thursday wasn't booked. -A lookup doesn't always need to bring the conversation to a stop. While it runs, the speaker can collect information that doesn't depend on the result. +### Questions and changes of topic -For example, if your service can retrieve a whole day's availability, the speaker can ask about morning or afternoon while that request is running. If the lookup requires that preference as an input, ask first. +A detour can help the caller make the next decision: -```mermaid -sequenceDiagram - participant C as Caller - participant S as Speaker - participant R as Reasoner and tools - C->>S: What is available on Friday? - S->>R: Look up Friday availability - S->>C: Morning or afternoon? - C->>S: Afternoon, ideally after two - R-->>S: Available times - S->>C: Explain matching options -``` +> **Caller:** Before I pick, what happens if I need to change it later? +> +> **Assistant:** Let me check that. *(service information comes back)* You can change or cancel up to 24 hours ahead at no charge. Would you still like Friday? -This illustrates a conversation to design and test, not a guarantee of exact event timing. Your prompts and tool requirements determine which work can overlap. +If the answer is already known, as with an appointment's length in the [example prompt](#example-split-a-scheduling-prompt), the speaker can give it directly. If it needs a lookup, delegate it. Either way, return to the unfinished task with the caller's earlier details intact. -Keep updates proportionate. One short “I'm checking” can be enough. A forced acknowledgment before every tool call, repeated status updates, or unrelated small talk can make the wait feel longer. +### Start with one assistant -### Confirming results +GPT-Live keeps one voice for the whole call. Start with a clear task and the tools it needs. Put its procedure in the reasoner prompt and its delegation triggers in the speaker prompt. Expand the assistant's responsibilities after it handles that task well. -The assistant can acknowledge the request before the work finishes. Its confirmation should follow the result: +Existing squads and assistant handoffs aren't supported with GPT-Live. See [Migrate to GPT-Live](/gpt-live/migrate) for moving a single assistant and the limits for squads. -| Moment | What the caller can hear | -| --- | --- | -| The request is understood | “I'll check Friday afternoon.” | -| The lookup finishes | “There are openings at two and four.” | -| A booking tool confirms success | “Your two o'clock appointment is booked.” | +## Names, numbers, and dates -The [example assistant](/gpt-live/configuration#create-an-assistant) only looks up availability. Add a booking tool before teaching it to offer the final step. +The speaker hears the caller's voice, but the reasoner works from the transcript, and a transcript can contain mistakes. A name, reference number, or date that's slightly wrong in the transcript becomes a wrong tool argument. -### When callers change the request +For anything the task depends on, have the assistant read it back before acting: -If the caller says “Actually, Thursday” during the lookup, the speaker should acknowledge the new request. Your application should match each result to the request it answers so Friday's result doesn't drive Thursday's booking. +> **Caller:** It's under Siobhan. +> +> **Assistant:** Could you spell that for me? +> +> **Caller:** S-I-O-B-H-A-N. +> +> **Assistant:** Thanks. Friday at 2:30 downtown, under Siobhan. Shall I book it? + +Read back values when they matter, not after every sentence. A confirmation before a booking is worth the extra few seconds. Repeating back an ordinary preference usually isn't. -Interrupting speech doesn't cancel an action already sent to a service. Keep changes and cancellations explicit in your application. The [limitations FAQ](/gpt-live/limitations#does-an-interruption-cancel-a-tool-call) explains this boundary. +Times need the same care when callers and your business can be in different time zones. The example business works in one time zone, and its tools return times in that zone. If yours doesn't, have the assistant confirm which time zone a time is in, and pass the zone to your tool explicitly, for example as an IANA name such as `America/New_York`. ## Speaking style -Choose both a voice and a way of speaking. The voice establishes a starting sound. Speaker instructions shape pacing, warmth, emphasis, and how much the assistant says at once. +The voice and the way it speaks are part of how well the call works. Decide based on who calls and why. -Write a short voice brief alongside the task brief. Describe behavior someone could hear and evaluate: +Compare two callers booking the same appointment. The first is booking for the first time and working through the details: -```text title="Example voice brief" -Sound warm and conversational. Keep routine answers brief. -Move briskly through familiar information, but slow down when reading -appointment dates and times. Give one instruction at a time when -explaining a process. Follow requests to speed up, slow down, or -explain more. Leave room for the caller to respond. -``` +> **Caller:** I've never been before, so I'm not sure what I need. +> +> **Assistant:** That's fine, we can go through it. A consultation is thirty minutes and doesn't need anything from you in advance. Would you like one at Downtown or Riverside? + +The second knows exactly what they want: -These are prompted behaviors, not precise speed or pause controls. Test the spoken result. Your application can also send [live speaker guidance](/gpt-live/configuration#live-call-control) with `append-context`; it doesn't guarantee exact wording, speed, or pauses. +> **Caller:** Standard visit, Riverside, next Tuesday morning. +> +> **Assistant:** Checking next Tuesday morning at Riverside. Name for the booking? + +The first caller benefits from a calm pace, fuller sentences, and room to think. The second wants short questions and no explanation. The same assistant can serve both if the speaker prompt describes how to read the caller: explain more when they're unsure, and move quickly when they've given everything. + +### Listening + +Brief acknowledgments such as "mm-hm" while the caller speaks can make the assistant sound attentive. Too many sound like interruptions. Set a level in the speaker prompt and adjust after listening to calls: + +```text +Backchannel policy: Use light backchannels. Acknowledge briefly without +competing with the caller. + +Interruption policy: Stop speaking when the caller interrupts. Listen to +what they say. +``` ### Pacing -A consistent personality doesn't require a constant pace. The same assistant can move quickly through familiar information and become more measured when a caller needs help. +A consistent personality doesn't require a constant pace. -| Moment | Delivery to try | +| Moment | Delivery to aim for | | --- | --- | -| Caller knows what they want | Short questions and concise answers | -| Caller is choosing between options | A few options at a time, with room to respond | -| Assistant reads a date or unfamiliar name | Slower delivery and clear articulation | -| Caller asks “Can you walk me through it?” | One step at a time | -| Caller says “Just the short version” | A brief summary with an offer to explain more | +| The caller knows what they want | Short questions and concise answers | +| The caller is choosing between options | Two or three options at a time, with room to respond | +| Reading a date, time, or unfamiliar name | Slower, clear delivery | +| "Can you walk me through it?" | One step at a time | +| "Just the short version" | A brief summary, with an offer to explain more | -Treat those requests as part of the conversation design. Test “slow down,” “a little faster,” and “say that another way” alongside your business scenarios. +These are prompted behaviors. GPT-Live doesn't have fixed speed or pause settings, so listen to calls to check the result. Test "slow down", "a bit faster", and "say that another way" alongside your business scenarios. ### Personality packs -Vapi's personality packs add speaking-style guidance to your speaker prompt: +Personality packs add speaking-style guidance to the speaker prompt. Each encourages a different way of listening and delivering: | Pack | What its guidance encourages | | --- | --- | -| **Active listening** (`eager-listener`) | Brief acknowledgments at natural openings, with room for the caller to continue | -| **Humming** (`idle-hummer`) | Soft humming or wordless sounds during quiet moments while waiting for tools | -| **Bouncy** (`bouncy`) | More expressive pitch, emphasis, and rhythm, adapting to the caller's mood | -| **Unhurried** (`unhurried`) | A slower pace, clear articulation, and space between ideas | +| `eager-listener` | Brief acknowledgments at natural openings, with room for the caller to continue | +| `unhurried` | A slower pace, clear articulation, and space between ideas | +| `bouncy` | More expressive pitch, emphasis, and rhythm, adapting to the caller's mood | +| `idle-hummer` | Soft humming or wordless sounds during quiet moments while waiting for tools | + +Choose a pack for the call's purpose. `unhurried` suits the first-time caller above, and might slow down the caller who wants a quick booking. `idle-hummer` changes how a blocked wait sounds, which may suit an informal service and feel out of place on a serious support call. Expressiveness and speed are separate choices: an animated voice can still speak slowly. + +To evaluate a pack, run the same scenario with and without it and listen to both. Packs shape delivery. They don't guarantee exact pauses, and they don't fix missed delegation or unclear task instructions. Field details are in [Settings and compatibility](/gpt-live/configuration#personality-and-language). + +### Language and pronunciation + +Write the speaker prompt and first message in the language you want the assistant to speak. Choose a voice by listening to it in that language with your own content; a voice's regional tag describes its accent, not the full set of languages it handles well. For names or terms that are often mispronounced, give a pronunciation cue in the speaker prompt, and ask the caller to spell names that are unclear. See the [voice list](/gpt-live/configuration#voices) for samples. + +## When the caller goes quiet + +A quiet caller and a pending result are different situations. When the assistant is waiting on your service, the design choices in [While work is running](#while-work-is-running) apply. When the caller has stopped responding, use idle-message hooks to check in. + +Configure them with `customer.speech.timeout` [assistant hooks](/assistants/assistant-hooks) and `say.prompt`. GPT-Live generates each check-in from your prompt, so the wording varies and isn't guaranteed. This example checks in twice, then ends the call if the caller still hasn't responded: + +```json title="Two check-ins and an optional final hangup" +{ + "hooks": [ + { + "on": "customer.speech.timeout", + "options": { + "timeoutSeconds": 8, + "triggerMaxCount": 3, + "triggerResetMode": "onUserSpeech" + }, + "do": [ + { + "type": "say", + "prompt": "Briefly check whether the caller is still there. If they said they needed a moment, let them know there's no rush." + } + ] + }, + { + "on": "customer.speech.timeout", + "options": { + "timeoutSeconds": 16, + "triggerMaxCount": 3, + "triggerResetMode": "onUserSpeech" + }, + "do": [ + { "type": "say", "prompt": "Briefly ask whether the caller would like more time." } + ] + }, + { + "on": "customer.speech.timeout", + "options": { + "timeoutSeconds": 30, + "triggerMaxCount": 1 + }, + "do": [ + { "type": "say", "prompt": "Say a brief goodbye because the caller hasn't responded." }, + { "type": "tool", "tool": { "type": "endCall" } } + ] + } + ] +} +``` -Choose a pack because it suits the experience. Humming might fit an informal concierge and feel out of place during a serious support issue. Expressiveness and speed are separate choices: an animated assistant can still speak slowly and clearly. +Here's what the caller experiences with this configuration: -Start with one pack, listen, and adjust the speaker prompt. Packs add guidance. They don't guarantee identical delivery on every call. See [personality configuration](/gpt-live/configuration#personality-and-language) and [voice previews](/gpt-live/configuration#voices). +- **Check-ins.** A quiet period starts after the assistant speaks. After 8 seconds without a response, the assistant checks in. At 16 seconds it offers more time. Both delays count from the same quiet period, so the first check-in doesn't push back the second. +- **The caller returns.** When the caller speaks, the check-ins stop. The next quiet period starts after the assistant's next reply. A check-in that was already being spoken may still be heard. +- **The caller asked for a moment.** Hooks can't tell a caller who is looking for their calendar from one who has left. If the assistant replies "Take your time", that reply starts a new quiet period, so the first check-in follows 8 seconds later. The first prompt asks the assistant to take the request into account, and longer timeouts give callers more room. +- **Final hangup.** At 30 seconds, the third hook asks for a goodbye and then ends the call. Once this hook fires, the call ends even if the caller starts speaking, so choose the timeout with care. The goodbye may be cut short. -## Testing conversations +Background noise doesn't count as the caller responding. Activity is measured from speech that appears in the transcript. + +Check-ins wait while the reasoner is actively working. An [async tool](/gpt-live/tools#slow-and-external-work) is different: once it has been dispatched, check-ins can resume even though its result is still pending. If callers may sit through a long external job, tell them what's happening, and set timeouts that leave room for it. + +Idle hooks are for a silent caller. They don't fix an assistant that promised to do something and never delegated it. For that, see [When to delegate](#when-to-delegate). Field details are in [Settings and compatibility](/gpt-live/configuration#idle-messages). + +## Turn the design into prompts + +Each decision on this page becomes a line in one of two prompts. + +### Where should an instruction go? -Review a short set of conversations before expanding the prompt. Include a direct request, a detour, information supplied out of order, a slow lookup, and a request to change speaking style. +Use `model.speaker.instructions` for **how to conduct the conversation and when to ask for help**. Use `model.reasoner.instructions` for **how to decide and act once asked**. -Ask whether the assistant used what the caller had already said, delegated at the right moment, and returned naturally to unfinished work. Listen for whether its pace helps the caller make progress. +| Instruction | Where it belongs | Why | +| --- | --- | --- | +| "Ask one question at a time and speak briefly." | Speaker | Shapes what the caller hears | +| "Delegate availability checks once you know the service, location, and date." | Speaker | Tells it when to start work | +| "Collect the booking name while availability is checked." | Speaker | Uses the wait | +| "Call `bookAppointment` only for a time from the latest lookup." | Reasoner | A condition for taking an action | +| "Never say a booking succeeded before the tool confirms it." | Both, worded for each | The reasoner must verify it and the speaker must describe it accurately | -Keep a few representative calls as a baseline. Repeat them when either prompt changes. Use [post-call analysis](/gpt-live/configuration#review-and-monitor-calls) to track outcomes, while reviewing audio for timing and delivery. Use [Voice Simulations](/gpt-live/limitations#can-i-run-simulations-and-evals) to repeat scenarios with GPT-Live as the tester, target, or both. Choose Voice mode and a single assistant target. +The prompts are separate. The reasoner doesn't see the speaker prompt, and the speaker doesn't see the reasoner prompt. Put rules that affect both the conversation and the actions in both, worded for each role. Keep detailed procedures in the reasoner prompt instead of copying one prompt into both fields. + +Your service remains responsible for enforcing rules that matter. A prompt that says "only book confirmed slots" guides the model. Your booking handler checking that the slot is still open is what makes it true. + +### Example: split a scheduling prompt + +These are the prompts for the appointment assistant used throughout this page. Its tools are described in [Build with tools](/gpt-live/tools). The speaker prompt follows a structure with labeled policy sections, which makes each decision easy to find and adjust: + +```text title="model.speaker.instructions" +You are the booking assistant for Example Service Studio, a fictional business +used for testing. Help callers book or cancel an appointment and answer +questions about the services. The call is done when the caller has a confirmed +booking they understand, has decided not to book, or has had their question +answered. + +Speak warmly and plainly. Keep most replies to one or two sentences. Ask one +question at a time. Use details the caller has already given and don't ask for +them again. Slow down when you say dates, times, and names. + +Facts you can share without checking: +- Callers should arrive five minutes early. +- A consultation is 30 minutes. Nothing is needed in advance. Callers can + bring any questions they want to cover. +- A standard visit is an hour. Callers can bring the reference number from a + previous visit, if they have one. + +Backchannel policy: Use light backchannels. Acknowledge briefly without +competing with the caller. + +Interruption policy: Stop speaking when the caller interrupts. Listen to what +they say. + +Delegation policy: +Backend tools: +- Appointments: check open times, book a time, and cancel a booking made on + this call. +- Service information: services, durations, what to bring, locations, hours, + and the change policy. +- Ending the call. + +Delegate to the backend when: +- You know the service, location, and date. Ask for availability right away. +- The caller has clearly said yes to booking after you read back the day, + time, location, and name. +- The caller changes the service, location, date, or chosen time. +- The caller asks to cancel or move a booking. +- The caller asks about services, what to bring, locations, hours, or policies. +- The backend asked for a detail and the caller has now given it. +- The caller asks to end the call or says goodbye. Always delegate this. + Saying goodbye does not end the call. + +Do not delegate to the backend when: +- You can answer from what was already said or from a result that still + answers the current request. +- You need a brief clarification first. + +While work is running: +- If there is a useful question that doesn't change the request, ask it. + Collect the name for the booking while availability is checked. +- If the caller asks something you can answer from information already given, + answer it. +- Otherwise acknowledge the wait once and give the caller room. Don't repeat + that you are still checking and don't guess how long it will take. + +Results: +- Say a time is booked only after the backend reports that it is booked. +- Offer only times from the latest lookup for the caller's current request. +- If the caller changed the request while work was running, delegate the + updated request and don't present the old results. +- When the caller chooses a time, that's a choice, not agreement to book. + Read back the day, time, location, and name, and ask whether to book it. + Delegate the booking only after the caller clearly says yes. +``` + +```text title="model.reasoner.instructions" +You handle appointment work for Example Service Studio, a fictional test +business, using the tools provided. Work from the latest request in the +conversation transcript. Transcripts can contain mistakes and later +corrections. + +Dates: If you don't know today's date, call getServiceInfo, which reports it +along with the bookable date range. Resolve relative dates such as "Friday" +against it. If a date is still ambiguous, return a short question for the +assistant to ask. + +Availability: Call lookupAvailability when you have the service (consultation +or standard-visit), the location (downtown or riverside), and a date as +YYYY-MM-DD. If a detail is missing, return a short question for the assistant +to ask. The lookup returns every open time for that day and location, so you +can filter by a time-of-day preference without another lookup. A different +service, location, or date needs a new lookup. + +Booking: A caller choosing a time is a selection, not agreement to book. Call +bookAppointment only when the transcript shows that, after the caller chose a +time from the latest lookup, the assistant read back the day, time, location, +and name and asked whether to book, and the caller then clearly said yes. If +that read-back and clear yes aren't in the transcript, don't book. Instead, +return the day, time, location, and name for the assistant to read back, and +say that it needs the caller's confirmation. Use that time's slotId. If the +caller changed the date, location, or service after that lookup, look up again +first. + +Changes: To move a booking made on this call, first look up the new time and +confirm with the caller that they want to move to it. Then cancel the existing +booking with cancelAppointment and book the new time. These are two separate +steps. If the new booking fails after the cancellation, say that the original +booking was cancelled and the new time wasn't booked, and offer to look again. + +Service questions: Use getServiceInfo. + +Results: Return short, plain facts the assistant can say: the open times, the +booking status and booking ID, and anything the caller needs to know next. No +Markdown or lists. When a tool fails, say what failed and the next step. Never +report a booking or cancellation that a tool did not confirm. If an outcome is +unclear, say so instead of guessing. + +Ending: Call endCall when the caller asks to end the call or says goodbye. +``` + +How the sections map to the decisions on this page: + +- **The first paragraph** states the caller's aim and what a finished call means, so the speaker knows when it's done. +- **Facts you can share without checking** gives the speaker stable answers it can use during a wait, without a lookup. +- **The delegation policy** lists every trigger, including answers to the reasoner's questions and ending the call. +- **While work is running** covers both a wait that can be used and one that's blocked, and rules out filler. +- **Results** covers stale results, corrections, and read-back before an action. +- **The reasoner prompt** holds the procedures and conditions for each tool, and asks for short, speakable results. The reasoner's output reaches the speaker, so Markdown or long explanations make the speaker's job harder. + +Treat these as a starting point. Change one thing at a time, and listen to the same scenarios after each change. + +The speaker prompt's policy headings follow the structure in OpenAI's [GPT-Live prompting guide](https://developers.openai.com/api/docs/guides/live-prompting), which has more model-level examples. When you apply its guidance, use Vapi's fields: conversation guidance goes in `model.speaker.instructions` and task procedures in `model.reasoner.instructions`. + +## Testing conversations -For additional model-level examples, see OpenAI's [GPT-Live prompting guide](https://developers.openai.com/api/docs/guides/live-prompting). Use Vapi's fields when applying that guidance to your assistant. +Listen to calls as well as reading transcripts: timing, pace, and how the assistant handles overlap only show up in audio. [Test and improve](/gpt-live/testing) has a scenario set built around the patterns on this page, including the wait patterns, corrections, and detours, along with ways to repeat them with Voice Simulations. diff --git a/fern/gpt-live/limitations.mdx b/fern/gpt-live/limitations.mdx index be24ad207..61dde8420 100644 --- a/fern/gpt-live/limitations.mdx +++ b/fern/gpt-live/limitations.mdx @@ -1,153 +1,56 @@ --- title: GPT-Live limitations and FAQs -subtitle: Check compatibility before you build or migrate -description: Check GPT-Live support for squads, call control, commentary injection, transfers, simulations, voices, fallbacks, recordings, and post-call data in Vapi. +subtitle: This page has moved +description: The GPT-Live limitations and FAQs now live in Settings and compatibility and the other GPT-Live guides. Each former section links to its new home. slug: gpt-live/limitations --- -GPT-Live has a different set of supported features from Vapi's other voice architectures. Use this page to check the requirements of your application before migrating. - -These limits describe **Vapi's GPT-Live integration**. A capability in OpenAI's direct API isn't necessarily exposed through Vapi. - -**Jump to:** [Calls and voices](#calls-and-voice) · [Transfers](#transfers-and-keypad-input) · [Squads](#can-i-use-an-existing-squad) · [Live control](#can-i-inject-commentary-or-control-speech-during-a-call) · [Idle messages](#can-i-use-idle-message-hooks) · [Testing](#can-i-run-simulations-and-evals) · [Recordings](#recordings-and-post-call-data) +The content of this page has moved. Supported features are in [Settings and compatibility](/gpt-live/configuration#compatibility), and the answers to these questions are in the guides linked below. ## Calls and voice -| Feature | Support | -| --- | --- | -| Browser WebRTC, native Twilio, and Vapi SIP | Supported | -| Raw WebSocket | Supported with mono, signed 16-bit little-endian PCM at 24 kHz. Set the rate explicitly | -| Native Telnyx or Vonage | Unsupported | -| 22 preset OpenAI voices | Supported. See [voice previews](/gpt-live/configuration#voices) | -| Custom or cloned voices, or other voice providers | Unsupported | -| Changing voice during a call | Unsupported. Save the assistant and start a new call | -| Model and voice fallbacks | Unsupported. GPT-Live also cannot be a fallback model | -| Reconnecting to an existing GPT-Live session | Unsupported. Start a new call and check the status of pending work | - -See [Connect a call](/gpt-live/configuration#connect-a-call) for working connection examples. +See [Calls and voice](/gpt-live/configuration#calls-and-voice) in Settings and compatibility. ## Transfers and keypad input -Function tools, API request tools, MCP tools, and `endCall` are supported. MCP tools run through the reasoner; see [MCP configuration](/gpt-live/configuration#mcp-tools). The other supported tool types, `transferCall` and `dtmf`, depend on the call connection: - -| Feature | Support | -| --- | --- | -| Blind transfer | Supported on native Twilio and Vapi SIP to phone or SIP destinations. Configure destinations for a `transferCall` tool, or supply one in an HTTP `transfer` request | -| Transfer on browser or raw WebSocket calls | Unsupported | -| Warm transfer or generated transfer summary | Unsupported | -| Transfer fallback plans or SIP `bye` verb | Unsupported | -| SIP transfer with extension dialing | Unsupported. Use a direct destination | -| Outgoing DTMF | Supported on SIP using RTP DTMF. Pass digits as a string, for example `"0011*#"` | -| Outgoing DTMF on Twilio, browser, or raw WebSocket calls | Unsupported | -| SIP INFO DTMF | Unsupported | -| Incoming keypad collection | Unsupported. Leave `keypadInputPlan.enabled` disabled | - -Other tool types are unsupported. Expose external lookups through a function, API request, or MCP tool. Carrier acceptance of a transfer doesn't establish that the destination answered. +See [Transfers and keypad input](/gpt-live/configuration#transfers-and-keypad-input) in Settings and compatibility, and [Transfer to a person](/gpt-live/tools#transfer-to-a-person). ## Can I use an existing squad? -Squads and assistant handoffs are currently unsupported. GPT-Live's speaker and reasoner work within one assistant. Delegation doesn't switch the caller to another assistant. - -Early testing suggests some use cases no longer need the same divisions between assistants. See [Start with one assistant](/gpt-live/design#start-with-one-assistant) for how to rethink the architecture while keeping your business requirements. +Squads and handoffs aren't supported with GPT-Live. See [Migrate to GPT-Live: From a squad](/gpt-live/migrate#from-a-squad). ## Can I inject commentary or control speech during a call? -Yes. Send HTTP requests to the call's `monitor.controlUrl` to: - -- End the call with `end-call`. -- Append `commentary`, `thinking`, or `instructions` to the speaker with `append-context`. -- Initiate a cold transfer with `transfer` on native Twilio or Vapi SIP calls. - -See [Live call control](/gpt-live/configuration#live-call-control) for request examples, authentication, and response behavior. Appended context affects the speaker; it doesn't update reasoner history or directly trigger tools. Submission doesn't guarantee exact wording or completed speech. - -The raw WebSocket connection supports audio streaming only. Send controls over HTTP, not WebSocket text frames or SDK data channels. Classic HTTP `say`, `add-message`, and mute/unmute commands remain unsupported. The `say` and `message.add` actions in [idle-message hooks](#can-i-use-idle-message-hooks) are separate assistant configuration. - -A caller can ask the assistant to slow down or explain differently, and you can prompt it to respond to those requests. That conversational behavior is separate from an application sending a control command. +See [Live call control](/gpt-live/configuration#live-call-control). ## Can I use idle-message hooks? -Yes. GPT-Live supports `customer.speech.timeout` hooks with `say.prompt`, `message.add`, and inline `endCall` actions. Idle messages are generated from the prompt; exact wording isn't guaranteed. Use separate hooks with staggered timeouts for multiple check-ins. Existing trigger limits and reset modes apply; a check-in doesn't postpone later hooks in the same silence window. - -Caller speech cancels remaining check-ins for that window. A `say` action requests one response and instructs the speaker not to repeat or resume it after an interruption. A hook's explicit `endCall` is different: once it fires, caller speech doesn't cancel the pending hangup. - -Saved tool references, function calls, transfers, and other actions aren't supported by these hooks. Support for idle messages doesn't imply support for every assistant hook event. See [Idle messages](/gpt-live/configuration#idle-messages) for examples and supported actions. Hook speech is generated and remains subject to the wording limitations below. +Yes. See [When the caller goes quiet](/gpt-live/design#when-the-caller-goes-quiet) for an example and [Idle messages](/gpt-live/configuration#idle-messages) for the settings. ## Can I guarantee exact speech or a fixed pause? -GPT-Live generates speech and may paraphrase supplied text. Speaker instructions and personality packs guide delivery. They don't provide exact playback, a fixed speaking rate, or a programmatic pause while work completes. - -Audio URL greetings and the generated-message first-message mode are unsupported. Use a text `firstMessage` with `assistant-speaks-first`, or let the assistant wait for the caller. Without usable greeting text, it waits. - -If exact prerecorded wording or strict control of every spoken step is essential, choose an architecture with those controls. See [voice design](/gpt-live/design#speaking-style) for the choices available through prompting. +No. GPT-Live generates all of its speech. See [Speaking style](/gpt-live/design#speaking-style). ## Does the reasoner get the conversation transcript? -Yes. Each delegation receives the available caller and assistant transcript up to that point, along with the reasoner's instructions, tool definitions, and retained reasoner/tool history. It isn't limited to a short snippet chosen by the speaker. - -It doesn't receive raw audio or automatically inherit the speaker prompt. The transcript is a snapshot at delegation time; later speech doesn't automatically update a delegation already in progress. See [reasoner context](/gpt-live/design#what-context-does-the-reasoner-receive) for the context comparison and correction example, and [prompt guidance](/gpt-live/design#where-should-an-instruction-go) for what to put in each prompt. +See [What context does the reasoner receive?](/gpt-live/design#what-context-does-the-reasoner-receive) ## Which existing settings no longer apply? -Classic transcriber selection, text-to-speech controls, and start/stop-speaking plans don't configure the GPT-Live conversation. General model fields such as `temperature` and `maxTokens` don't tune this integration. - -Use the speaker prompt for conversation behavior and the reasoner settings for task reasoning. The reasoner currently supports only the OpenAI models listed in [Model settings](/gpt-live/configuration#model-settings). +See [Settings that no longer apply](/gpt-live/configuration#settings-that-no-longer-apply) and the [settings map for migration](/gpt-live/migrate#4-map-the-settings). ## Does an interruption cancel a tool call? -No. Interrupting speech changes the conversation, not an action already submitted to your service. Your application needs to check whether that action is pending, completed, or cancellable. - -If the caller corrects a date during a lookup, use the result that matches the updated request. For an example, see [When callers change the request](/gpt-live/design#when-callers-change-the-request). +No. See [When callers change the request](/gpt-live/design#when-callers-change-the-request). ## Can I run Simulations and Evals? -Yes. **Voice Simulations support GPT-Live** as the assistant under test, the AI tester, or both. You can also pair a GPT-Live participant with a classic voice assistant. - -| Simulation setup | Support | -| --- | --- | -| GPT-Live target with a classic AI tester | Supported in Voice mode | -| GPT-Live AI tester with a classic assistant target | Supported in Voice mode | -| GPT-Live tester and target | Supported in Voice mode | -| Chat mode with either participant using GPT-Live | Unsupported; choose Voice | -| Squad runs with a GPT-Live participant | Unsupported; choose a single assistant target | - -To test your assistant: - -1. Follow the [Simulations quickstart](/observability/simulations-quickstart), selecting your GPT-Live assistant as the target. -2. Configure the scenario, AI tester personality, and success criteria. If you use a GPT-Live tester, select a supported native OpenAI voice. -3. Run the suite in **Voice** mode, then review its transcript, evaluations, and available recording. - -Existing GPT-Live organization access requirements still apply. Simulations handle the participants' audio formats, including mixed GPT-Live and classic runs; you don't need to build an audio bridge yourself. [Tool mocks](/observability/simulations-advanced#mock-tool-responses) are supported for GPT-Live simulation tools. Unmocked tools can perform real actions, so configure mocks or sandbox services before running scenarios that change external data. - -Voice Simulations let you exercise speech, timing, interruptions, and tool behavior. Listen to the recording as well as checking evaluations: a passing criterion doesn't establish every aspect of conversation quality. [Evals](/observability/evals-quickstart) remain useful for supported model decisions, but don't replace testing the voice interaction. +See [Repeat scenarios with Voice Simulations](/gpt-live/testing#repeat-scenarios-with-voice-simulations). ## Recordings and post-call data -GPT-Live supports post-call analysis, structured outputs, scorecards, Boards, and Monitoring. Availability depends on your configuration and retained call data: - -| Setting or feature | Behavior | -| --- | --- | -| Recording disabled in the assistant | No recording is created | -| Recording-consent plan configured | Recording is disabled because GPT-Live doesn't collect that consent | -| Transcript-based analysis | Requires the relevant conversation data | -| Audio-based analysis | Requires an available recording | -| Monitoring with Zero Data Retention | Unsupported | -| Live listen or monitor sockets | Unsupported. Post-call Monitoring doesn't provide a live audio connection | - -See [Review and monitor calls](/gpt-live/configuration#review-and-monitor-calls) for setup and result fields. +See [Recordings and post-call data](/gpt-live/configuration#recordings-and-post-call-data). ## Troubleshooting -Use the symptom to find the setting or handler to check: - -| Symptom | Check | Fix | -| --- | --- | --- | -| GPT-Live is unavailable | Organization feature access | Contact Vapi to enable access. | -| Voice unsupported | `voice.provider` and `voice.voiceId` | Use OpenAI and a [supported voice ID](/gpt-live/configuration#voices). | -| Tool configuration unsupported | Tool type and resolved names | Use supported types and unique names. Inspect saved tools too. | -| Call fails on WebSocket startup | Audio format and sample rate | Set raw `pcm_s16le` at `24000` Hz explicitly. | -| WebSocket audio sounds distorted | Sample encoding, rate, and channels | Convert the source to mono, signed 16-bit little-endian PCM at 24 kHz. | -| Assistant never greets | First-message mode and usable text | Set `assistant-speaks-first` and a text `firstMessage`. | -| Assistant promises a lookup but no tool runs | Both prompts and attached tools | Add a concrete delegation rule and make the tool available to the reasoner. | -| Tool result never reaches the conversation | Handler response and matching ID | Return the actual `toolCallId`. Return the final async result through the original handler request. | -| Transfer fails | Transport and configured destination | Use a supported blind transfer. Check carrier requirements and destination restrictions. | -| No recording | Recording settings and consent plan | Check the documented recording behavior before the test. | +See [Troubleshooting](/gpt-live/testing#troubleshooting) in Test and improve. diff --git a/fern/gpt-live/migrate.mdx b/fern/gpt-live/migrate.mdx new file mode 100644 index 000000000..01722eb78 --- /dev/null +++ b/fern/gpt-live/migrate.mdx @@ -0,0 +1,205 @@ +--- +title: Migrate to GPT-Live +subtitle: Move an existing assistant, compare it with the original, and keep a way back +description: Migrate a Vapi assistant to GPT-Live. Check fit, convert a copy, split prompts between the speaker and reasoner, compare results, and roll back safely. +slug: gpt-live/migrate +--- + +Move a copy of your assistant first, compare it with the original, and switch traffic after it meets the same goals. This guide covers a single assistant. If you use a squad, see [From a squad](#from-a-squad) for the current limits. + +## Before you start + +### Check that your assistant can move + +Some features don't carry over. If your assistant depends on one of these, plan how to replace it before you begin: + +- **Voices** must be one of the [OpenAI voices GPT-Live supports](/gpt-live/configuration#voices). Other voice providers and custom or cloned voices aren't available. +- **Squads and assistant handoffs** aren't supported. There is no direct squad conversion; see [the current limits](#from-a-squad). +- **Model and voice fallbacks** must be removed. GPT-Live can't be a fallback model either. +- **Exact or prerecorded speech** isn't available. GPT-Live generates all of its speech, including the greeting. +- **Keypad input** from callers isn't supported. +- **Knowledge bases and the Query tool** aren't available. Use a function, API request, or MCP tool for retrieval. +- **Transfers** are blind only, on native Twilio and Vapi SIP calls. **Native Telnyx and Vonage** numbers aren't supported. + +The full list is in [Settings and compatibility](/gpt-live/configuration#compatibility). + +### Save representative calls + +Pick a handful of calls that represent what the assistant has to get right: a straightforward success, a caller who changes their mind, a failure your tools return, and anything specific to your business. For each, note the starting state, the expected tool actions, the expected outcome, and a recording from the current assistant. These are what you'll compare against. + +### Keep the original running + +Leave your phone numbers on the original assistant until the GPT-Live version has passed the same calls. Work on a copy, test it on its own, and switch traffic when you're ready. + +## From a single assistant + +### 1. Take inventory + +Write down what the assistant does and what it depends on: + +| Area | What to record | +| --- | --- | +| Caller's goal | What a successful call achieves | +| Prompt | The system prompt, and which parts are about conversation and which are procedures | +| Tools | Each tool, its type, and when it should be used | +| Voice | The current provider and voice, and what matters about it | +| Call connection | Browser, Twilio, SIP, WebSocket, or another provider | +| Hooks | Idle messages and any other hooks | +| Transfers | Destinations, and whether any are warm transfers | +| Analysis | Structured outputs, scorecards, and anything that reads the transcript or recording | + +### 2. Convert a copy + +Convert a copy, so the original keeps serving callers while you test: + + + +In [Assistants](https://dashboard.vapi.ai/assistants), open the assistant's menu and select **Duplicate**. Don't attach a phone number to the copy. + + +Open the copy, select **Try GPT Live**, then confirm with **Switch to GPT Live**. + + +Conversion publishes a new version of the copy straight away, with speaker and reasoner prompts written for you. Read both prompts and compare them with the split described in the next step. + + + +Because conversion publishes immediately, working on a duplicate is what keeps callers on the original until you're ready. + +To go back, select **Revert to Classic**. It restores the assistant's Classic configuration as a draft, and the published version stays GPT-Live until you select **Publish**. + +You can also migrate by hand through the API: create a new assistant with the GPT-Live configuration from the steps below. + +### 3. Split the prompt + +A classic assistant has one prompt that covers both how to talk and how to do the work. GPT-Live has two. The speaker prompt shapes the conversation and says when to delegate. The reasoner prompt holds the procedures and tool rules. Neither sees the other, so rules that affect both belong in both, worded for each role. + +For example, a classic prompt for an appointment assistant at Example Service Studio, a fictional business, might read: + +```text title="Before: one classic system prompt" +You are the booking assistant for Example Service Studio. Be friendly and brief. +Ask for the service, location, and date. Use lookupAvailability to find open +times and read them to the caller. When they choose one, ask for their name, +confirm the details, and call bookAppointment. Never say a booking is done +until bookAppointment succeeds. If the caller wants to change a booking, cancel +it with cancelAppointment first. Answer questions about services with +getServiceInfo. +``` + +Split it by who does what: + +```text title="After: speaker instructions" +You are the booking assistant for Example Service Studio. Speak warmly and +briefly, one question at a time. Use details the caller has already given. + +Delegation policy: +Backend tools: +- Appointments: check open times, book a time, and cancel a booking. +- Service information: services, what to bring, locations, hours, policies. +- Ending the call. + +Delegate to the backend when: +- You know the service, location, and date. Ask for availability right away. +- The caller has clearly said yes to booking after you read back the day, + time, location, and name. +- The caller changes the request, or asks to cancel or move a booking. +- The caller asks about services, what to bring, locations, hours, or policies. +- The backend asked for a detail and the caller has now given it. +- The caller asks to end the call or says goodbye. Always delegate this. + +While work is running, ask for the name for the booking if you don't have it. +If the caller asks something you can answer from what you already know, answer +it. Otherwise acknowledge the wait once and give the caller room. + +Say a time is booked only after the backend reports that it is booked. When +the caller chooses a time, that's a choice, not agreement to book. Read back +the day, time, location, and name, and ask whether to book it. Delegate the +booking only after the caller clearly says yes. +``` + +```text title="After: reasoner instructions" +Work from the latest request in the conversation transcript. Call +lookupAvailability when you have the service, location, and date. A caller +choosing a time isn't agreement to book. Call bookAppointment only for a time +from the latest lookup, and only when the transcript shows the assistant read +back the day, time, location, and name and the caller then clearly said yes. +Otherwise, return those details for the assistant to read back and confirm. +To move a booking, look up the new time and confirm +it with the caller first, then cancel the existing booking with +cancelAppointment and book the new time. If the new booking fails after the +cancellation, say so. Use getServiceInfo for service questions. Return short, +plain facts. Never report an action a tool did not confirm. Call endCall when +the caller asks to end the call. +``` + +Three things changed besides the split: + +- **Explicit delegation triggers.** The speaker only starts work when it delegates, so every kind of task, including ending the call, is listed. +- **Waiting is designed.** In the classic flow, the assistant read times out after a lookup and then asked for the name. The speaker now asks for the name while the lookup runs. +- **Results are for speaking.** The reasoner returns short facts that the speaker can use directly. + +### 4. Map the settings + +| Classic setting | With GPT-Live | +| --- | --- | +| System prompt | Conversation guidance goes in `model.speaker.instructions`, procedures in `model.reasoner.instructions`. If speaker instructions are set, they take precedence over the classic system prompt | +| Model | `model.model` is `gpt-live-1`. The reasoner model is set separately and defaults to `gpt-5.6-terra` | +| Transcriber | Not used. GPT-Live hears the caller directly | +| Voice | An OpenAI voice from the [supported list](/gpt-live/configuration#voices) | +| Model or voice fallbacks | Remove them | +| Function, API request, and MCP tools | Keep them. Move their usage rules into the reasoner prompt | +| `endCall`, `transferCall`, `dtmf` | Keep them, and check [transfer and keypad support](/gpt-live/configuration#transfers-and-keypad-input) for your connection | +| Other tool types | Not supported. Replace them with a function, API request, or MCP tool | +| Tool messages, such as request-start messages | Not spoken. Guide what the assistant says while work runs in the speaker prompt | +| Idle-message hooks with `say.prompt` | Keep them. See [When the caller goes quiet](/gpt-live/design#when-the-caller-goes-quiet) | +| Idle-message hooks with `say.exact` | Accepted, but the text is spoken as generated speech, so the words may vary and verbatim playback isn't guaranteed. Rewrite them with `say.prompt` | +| Greeting | A text `firstMessage` works and is spoken as generated speech. Audio greetings and generated first-message mode aren't supported | +| Keypad input plan | Must be disabled | +| Start and stop speaking plans, `temperature`, `maxTokens` | No effect | +| Structured outputs, call analysis, scorecards | Keep them. See [Test and improve](/gpt-live/testing#watch-production-calls) | + +### 5. Redesign the waits + +Look at the places where your current flow waits for a tool before it moves on. With GPT-Live, the conversation can keep going during that time, so decide what it should cover. Start work as soon as its inputs are known, then ask for something the next step will need: + +> **Caller:** A consultation downtown on Friday. +> +> **Assistant:** I'll check Friday downtown. Who should I put the appointment under? +> +> **Caller:** Sam Lee. +> +> **Assistant:** Thanks, Sam. Friday I have 10:30, 2:30, or 4. + +When the next step is blocked, the conversation can still help. Here the booking is running, and the caller asks about something the assistant already knows: + +> **Assistant:** That's Friday at 2:30, downtown, under Sam Lee. Shall I book it? +> +> **Caller:** Yes. Do I need to bring anything? +> +> **Assistant:** Nothing's required for a consultation. Bring any questions you'd like to cover, and it helps to arrive five minutes early. +> +> **Assistant:** You're booked: Friday at 2:30 downtown, under Sam Lee. + +The preparation answer doesn't depend on the booking, so it can come while the booking runs. "You're booked" waits for the result. For the speaker to answer like this, give it a few stable facts in its prompt, such as arrival and preparation guidance. When there's nothing useful to say, one acknowledgment and a pause is better than filler. [Design conversations](/gpt-live/design#while-work-is-running) covers these choices in depth. + +### 6. Compare with the original + +Run your saved calls against the copy. For each one, check: + +- **The outcome.** Your service's state shows the same result as before, with no duplicate or missing actions. +- **The tool calls.** The call's messages show the expected tool calls and arguments. +- **The conversation.** The recording shows the assistant reached the goal without making the caller repeat themselves, and that waits sounded reasonable. + +Record differences you intended, such as a shorter call, separately from regressions. [Test and improve](/gpt-live/testing) describes how to repeat these calls with Voice Simulations. + +### 7. Roll out and roll back + +When the copy passes, assign your phone number to it. Keep the original assistant, unchanged, until you're confident in the new one. + +If you need to take a converted assistant back to Classic, select **Revert to Classic**, then **Publish**. Until you publish, callers still reach the GPT-Live version. + +## From a squad + +GPT-Live doesn't support squads or assistant handoffs. **Switch to GPT Live** converts one assistant; it doesn't convert a squad or preserve its routing and handoff behavior. + +Keep your existing squad if your call flow depends on those capabilities. The single-assistant steps above aren't a replacement for a squad migration. diff --git a/fern/gpt-live/overview.mdx b/fern/gpt-live/overview.mdx index 3cc4bae37..ac83e500b 100644 --- a/fern/gpt-live/overview.mdx +++ b/fern/gpt-live/overview.mdx @@ -1,72 +1,95 @@ --- title: GPT-Live -subtitle: Voice conversations with tools and reasoning built in -description: Learn what GPT-Live brings to Vapi, how its speaker and reasoner work together, and where to start with conversation design, configuration, and compatibility. +subtitle: Voice conversations that keep going while the work gets done +description: Learn what GPT-Live changes for a Vapi assistant, how its speaker and reasoner work together, whether it fits your use case, and where to start. slug: gpt-live/overview --- -GPT-Live is an OpenAI voice model that can listen and speak at the same time. A caller can clarify a request while the assistant is talking, ask a question during a lookup, or ask it to slow down. This is **full-duplex conversation**: audio can flow in both directions at once. +GPT-Live is an OpenAI voice model that listens to the caller's audio and generates its own speech. It keeps listening while it speaks, so it can respond to an interruption or a new detail while it's talking. -Vapi connects that conversation to your tools and phone numbers, and provides the call records, analysis, and monitoring around it. +In Vapi, a GPT-Live assistant pairs that voice model with a second model that does the task work, so the conversation can carry on while your tools run. -**Private beta.** GPT-Live must be enabled for your Vapi organization. If you don't have access yet, [join the waitlist](https://vapi.ai/gpt-live-waitlist). We're admitting users from the waitlist every day. +**Private beta.** GPT-Live must be enabled for your Vapi organization. See [Access](#access). If you already have a Vapi assistant, see [Migrate to GPT-Live](/gpt-live/migrate). -## In this guide +## How it works -Start here for an overview, then choose what you need next. +A GPT-Live assistant has two parts: - - - Plan speaker and reasoner roles, conversation flow, and speaking style. - - - Create an assistant, choose a voice, connect tools, and make a call. + + + The GPT-Live voice the caller hears. It listens throughout the call, responds to the caller, and decides when to ask the reasoner for help. - - Check supported features and what to consider before migrating. + + A separate model that does the task work. It follows your procedures, calls your tools, and returns what it found. -## How it works +When the speaker needs something done, such as checking a calendar, it **delegates** the task to the reasoner. The reasoner works from the conversation so far and your instructions. Meanwhile the speaker can continue listening and responding to the caller. When the result comes back, the speaker fits it into the current conversation. -A GPT-Live assistant uses a **speaker** for the spoken interaction and a **reasoner** for tasks that need tools or detailed procedures. **Delegation** is the speaker asking the reasoner to handle one of those tasks. +Because GPT-Live hears and speaks directly, you don't configure a separate transcriber or text-to-speech provider. -```mermaid -flowchart TD - C[Caller] <-->|Audio| S[Speaker] - S -->|Delegates a task| R[Reasoner] - R <-->|Requests and results| T[Tools and your services] - R -->|Result for the conversation| S -``` +## What changes with GPT-Live -For a scheduling assistant, the speaker can guide the caller toward a suitable appointment. The reasoner checks availability through your service and returns the options. The speaker can keep listening while that work happens. +For the caller, the conversation can feel less like taking turns. They can speak over the assistant, add a detail while it's talking, or ask something else while it's checking. -GPT-Live handles speech directly. You don't configure a separate speech-to-text transcriber or text-to-speech provider for the conversation. +For you, more of the design lives in the two prompts: -## What changes with GPT-Live +- **Decide what happens while work runs.** Which work to start early, what the assistant can usefully ask or answer while it waits, and how it brings the result back into the conversation. +- **Make delegation explicit.** The speaker only starts work when it delegates, so tell it which requests need the reasoner. +- **Plan for generated speech.** All speech, including the greeting, is generated, so exact wording isn't guaranteed. +- **Keep saying and doing separate.** The assistant saying something happened is different from your service confirming it. -You can design the interaction around the caller's goal, with room for questions and changes along the way. Detailed procedures belong with the reasoner. The speaker's instructions focus on how to guide the conversation. +[Design conversations](/gpt-live/design) works through these decisions with an appointment-booking example. -You also have more to design than the words alone. Set a voice, give the speaker a delivery brief, and use personality packs to explore different styles. A concise assistant can slow down for a date, explain one step at a time, or give the short version when asked. +## In this guide -[Design your assistant](/gpt-live/design) explains these choices with examples, including how to rethink an existing squad. + + + Create an assistant and talk to it in a few minutes, with no server of your own. + + + Shape a conversation that reaches the caller's goal while work runs alongside it. + + + Move an existing assistant, compare it with the original, and keep a way back. + + + See what's supported before you build or migrate. + + ## What Vapi provides -| Capability | Where to learn more | +Vapi connects GPT-Live to your phone numbers and tools, and provides call records and observability according to your recording and data settings: + +- Browser, Twilio, Vapi SIP, and WebSocket calls. +- Function, API request, and MCP tools, plus built-in end-call, transfer, and DTMF tools. +- Speaker and reasoner prompts, personality packs, and 22 voices. +- Idle-message hooks and HTTP live call control. +- Transcripts, recordings, structured outputs, scorecards, Boards, Monitoring, and Voice Simulations. + +Before you commit, check whether your assistant depends on something GPT-Live doesn't support: + +| If you need | With GPT-Live | | --- | --- | -| Browser, phone, and WebSocket calls | [Connect a call](/gpt-live/configuration#connect-a-call) | -| Function, API request, and MCP tools that look up information or take actions | [Connect your tools](/gpt-live/configuration#connect-your-tools) | -| HTTP live call control: end call, speaker context, and cold transfer | [Control an active call](/gpt-live/configuration#live-call-control) | -| Idle-message check-ins and optional call-ending hooks | [Configure idle messages](/gpt-live/configuration#idle-messages) | -| Speaker prompts, reasoner settings, and personality packs | [Model settings](/gpt-live/configuration#model-settings) | -| 22 preset voices, with audio previews | [Choose a voice](/gpt-live/configuration#voices) | -| Voice Simulations with GPT-Live testers and targets | [Test conversations](/gpt-live/limitations#can-i-run-simulations-and-evals) | -| Transcripts, available recordings, and post-call analysis | [Review calls](/gpt-live/configuration#review-and-monitor-calls) | -| Structured outputs, scorecards, Boards, and Monitoring | [Track quality](/gpt-live/configuration#review-and-monitor-calls) | - -GPT-Live has a different set of supported features from Vapi's other voice architectures. Check [Limitations and FAQs](/gpt-live/limitations) if your application depends on squads, live call control, simulations, or a specific transfer flow. - -For the underlying model, see OpenAI's [GPT-Live documentation](https://developers.openai.com/api/docs/guides/live). This section describes Vapi's integration and its supported configuration. +| Exact or prerecorded speech | Not available. All speech is generated | +| A voice from another provider, or a custom voice | Not available. Choose from 22 OpenAI voices | +| A squad with handoffs | Not supported. GPT-Live runs as a single assistant | +| Knowledge bases or the Query tool | Not available. Use a retrieval tool | +| Warm transfers, or transfers on browser calls | Not supported. Blind transfers work on Twilio and Vapi SIP | +| Caller keypad input | Not supported | +| Native Telnyx or Vonage numbers | Not supported | + +The full list is in [Settings and compatibility](/gpt-live/configuration#compatibility). + +These pages describe Vapi's GPT-Live integration. OpenAI's API and ChatGPT offer some capabilities, such as web search, custom voices, and image input, that aren't available through Vapi. For background on the model itself, see OpenAI's [GPT-Live documentation](https://developers.openai.com/api/docs/guides/live). + +## Access + +GPT-Live is in private beta and must be enabled for your organization. If you don't have access yet, [join the waitlist](https://vapi.ai/gpt-live-waitlist). Joining doesn't guarantee early access. + +## Cost + +A GPT-Live call is billed for voice time by the second, including silence and waiting, plus Vapi's platform fee, the reasoner's token usage, and telephony. See [Cost](/gpt-live/testing#cost) for a worked example and how to check a real call. diff --git a/fern/gpt-live/quickstart.mdx b/fern/gpt-live/quickstart.mdx new file mode 100644 index 000000000..432bb32d9 --- /dev/null +++ b/fern/gpt-live/quickstart.mdx @@ -0,0 +1,149 @@ +--- +title: GPT-Live quickstart +subtitle: Make a first call and see how the speaker and reasoner work together +description: Create a GPT-Live assistant in the Vapi dashboard or API with no custom backend, talk to it in the browser, and check the call record for the end-call action. +slug: gpt-live/quickstart +--- + +This quickstart creates a simple front-desk assistant and makes a browser call to it. It needs no server of your own. It uses one built-in tool, `endCall`, so you can see a delegation happen and find it in the call record. + +## Before you start + +- GPT-Live enabled for your Vapi organization. See [access](/gpt-live/overview#access). +- A browser with a microphone. +- For the API path, a Vapi private API key. +- If you use your own OpenAI API key, it needs access to GPT-Live and to the reasoner model. This quickstart uses the default, `gpt-5.6-terra`. + +## Create the assistant + +The assistant answers for Example Service Studio, a fictional business. It greets the caller, asks for their name and reason for calling, and ends the call when they're done. It has two prompts: the **speaker** prompt shapes the conversation, and the **reasoner** prompt handles tasks, which here is only ending the call. + + + + + +Open [Assistants](https://dashboard.vapi.ai/assistants) and select **Create Assistant**. This creates a new Classic assistant. + + +On the new assistant, select **Try GPT Live**, then confirm with **Switch to GPT Live**. The switch publishes the assistant, so don't attach a phone number until you've finished testing. When the conversion completes, choose **Marin** as the voice. + + +In **Speaker**, enter: + +```text title="Speaker instructions" +You are the front desk for Example Service Studio, a fictional business used +for testing. Greet callers warmly, ask for their name and the reason for their +call, and summarize what they told you. You can't look anything up or book +appointments yet, so say so if they ask. + +Speak plainly and keep replies short. Ask one question at a time. Use details +the caller has already given. + +Delegation policy: +Backend tools: +- Ending the call. + +Delegate to the backend when: +- The caller asks to end the call or says goodbye. Always delegate this. + Saying goodbye does not end the call. +``` + +In **Reasoner**, enter: + +```text title="Reasoner instructions" +Call endCall when the caller asks to end the call or says goodbye. +``` + + +Add the built-in **End Call** tool to the assistant. + + +Set the first message to "Hi, this is Example Service Studio. Who am I speaking with?" and set the assistant to speak first. Select **Publish**. + + + + + + +Save this as `quickstart-assistant.json`: + +```json title="quickstart-assistant.json" +{ + "name": "GPT-Live quickstart", + "model": { + "provider": "openai", + "model": "gpt-live-1", + "speaker": { + "instructions": "You are the front desk for Example Service Studio, a fictional business used for testing. Greet callers warmly, ask for their name and the reason for their call, and summarize what they told you. You can't look anything up or book appointments yet, so say so if they ask.\n\nSpeak plainly and keep replies short. Ask one question at a time. Use details the caller has already given.\n\nDelegation policy:\nBackend tools:\n- Ending the call.\n\nDelegate to the backend when:\n- The caller asks to end the call or says goodbye. Always delegate this. Saying goodbye does not end the call." + }, + "reasoner": { + "instructions": "Call endCall when the caller asks to end the call or says goodbye." + }, + "tools": [ + { "type": "endCall" } + ] + }, + "voice": { + "provider": "openai", + "voiceId": "marin" + }, + "firstMessageMode": "assistant-speaks-first", + "firstMessage": "Hi, this is Example Service Studio. Who am I speaking with?", + "maxDurationSeconds": 600 +} +``` + + +Set your Vapi private API key in the terminal, then create the assistant: + +```bash +export VAPI_PRIVATE_API_KEY="paste-your-private-key-here" + +curl --fail-with-body https://api.vapi.ai/assistant \ + -H "Authorization: Bearer $VAPI_PRIVATE_API_KEY" \ + -H "Content-Type: application/json" \ + --data-binary @quickstart-assistant.json +``` + +The reasoner uses its defaults: `gpt-5.6-terra` with `low` reasoning effort. Keep the returned `id` to open the assistant in the dashboard. + + + + + +## Talk to it + +Open the assistant in the dashboard and select **Talk**. Allow microphone access. Try these, one at a time: + +1. **Give several details at once.** "Hi, it's Sam, I'm calling about moving my appointment." The assistant should use both and not ask for your name again. +2. **Interrupt it.** Start speaking while it's talking. It should stop and listen. +3. **Change the pace.** "Could you slow down a little?" It should adjust. +4. **Say goodbye.** "Thanks, that's everything. Bye." The call should end. + +The greeting and replies are generated, so the wording varies from call to call and may not match your first message exactly. The assistant may also be cut off before it finishes saying goodbye, because ending the call is an action that happens as soon as the reasoner takes it. + +## Check what happened + +Open the call in [Call Logs](https://dashboard.vapi.ai/calls). You'll find: + +- **The transcript** of both sides of the conversation. +- **The recording**, if recording is enabled. +- **An `endCall` tool call** in the call's messages. This is the reasoner acting on the speaker's delegation. +- **The ended reason**, which shows that the assistant ended the call. + +That `endCall` entry is the thing to look for. The assistant saying "goodbye" and the call actually ending are separate events, so check for the action as well as the words. + +## If the call didn't end + +If the assistant said goodbye but the call stayed open, look in the call's messages for an `endCall` tool call: + +- **No `endCall` entry.** The speaker may not have delegated, or the reasoner may not have called the tool. Check that the speaker prompt lists ending the call as a delegation trigger, the reasoner prompt says to call `endCall`, and the tool is attached. Delegation is a model decision, so try a few calls. +- **An `endCall` entry is there.** The delegation happened, so look at the tool's result for an error. + +See [the troubleshooting guide](/gpt-live/testing#said-goodbye-but-the-call-didnt-end) for more checks. + +## Next steps + +- **[Design conversations](/gpt-live/design):** how to shape a conversation that reaches the caller's goal while work runs in the background. +- **[Build with tools](/gpt-live/tools):** connect your service so the assistant can look things up and take actions. +- **[Migrate to GPT-Live](/gpt-live/migrate):** move an existing assistant. diff --git a/fern/gpt-live/testing.mdx b/fern/gpt-live/testing.mdx new file mode 100644 index 000000000..820ec7f76 --- /dev/null +++ b/fern/gpt-live/testing.mdx @@ -0,0 +1,212 @@ +--- +title: Test and improve GPT-Live assistants +subtitle: Check the conversation and the outcome together, then fix the part that needs work +description: Test GPT-Live assistants with representative scenarios, Voice Simulations, and call records. Measure latency and cost, monitor production calls, and troubleshoot common symptoms. +slug: gpt-live/testing +--- + +A GPT-Live call can go wrong in two separate ways. The conversation can be poor even when the task succeeds, for example when the caller has to repeat themselves. And the task can fail even when the conversation sounds fine, for example when the assistant says "You're booked" and nothing was booked. Test both, using evidence for each. + +## What to look at + +| Evidence | What it tells you | +| --- | --- | +| **Recording** | What the caller heard, and when: pace, overlap, interruptions, how waits sounded | +| **Transcript** | What was said. Transcripts can contain recognition mistakes, so check key details against the recording | +| **Call messages** | Tool calls with their arguments and results, including `endCall` | +| **Your service's state** | What actually happened: bookings made, changed, or not made | +| **Ended reason** | How the call ended | +| **Costs** | The call's cost breakdown, including billable voice time | + +Tool results don't prove what the caller heard, and a good-sounding call doesn't prove the action happened. When you judge a scenario, check the one that answers your question: the recording for delivery, the call messages for what the reasoner did, and your service for the outcome. + +## Scenarios to run + +Start with a small set that covers what your assistant must get right, and repeat it after every prompt or setting change. Voice models vary from call to call, so run each important scenario several times. + +| Scenario | What to do | Passes when | +| --- | --- | --- | +| Direct request | Give everything the task needs at once | The assistant starts work without re-asking, and the outcome is correct | +| Missing detail | Leave out something required | It asks for that one detail, then continues | +| Useful wait | Ask for something that needs a lookup | It asks a relevant question while the lookup runs, and uses the answer | +| Blocked wait | Confirm an action and stay quiet | One acknowledgment, no invented progress, and the outcome is stated only after the result | +| Question during a wait | Ask something related while an action runs | It answers from known information and returns to the result | +| Correction | Change a detail while work runs | New work for the new detail. The old result isn't presented | +| Change after completion | Change your mind after an action succeeds | It says the action was done and follows your change procedure | +| Detour and return | Ask an unrelated question mid-task | It answers, then returns to the task without repeating completed work | +| Tool failure | Make a tool fail | It explains what failed. Nothing is claimed | +| Interruption | Talk over the assistant | It stops and responds to you. Running work continues, and your service's state is correct | +| Quiet caller | Stop responding | Check-ins as configured. They stop when you speak again | +| End the call | Say goodbye | The call ends through `endCall` | +| Transfer | Ask for a person, on a phone call | The transfer is requested and connects | + +To test failures and slow tools on demand, point your tools at a test version of your service that can return errors or delay its responses. [Actions that change state](/gpt-live/tools#actions-that-change-state) covers what your service should check in each case. + +## Repeat scenarios with Voice Simulations + +[Voice Simulations](/observability/simulations-quickstart) run scripted scenarios against your assistant with an AI tester, so you can repeat the set above without calling in yourself. GPT-Live can be the assistant under test, the tester, or both. + +| Setup | Support | +| --- | --- | +| GPT-Live assistant with a classic AI tester | Supported in Voice mode | +| GPT-Live AI tester with a classic assistant | Supported in Voice mode | +| GPT-Live tester and assistant | Supported in Voice mode | +| Chat mode with a GPT-Live participant | Not supported. Use Voice mode | +| Squad runs with a GPT-Live participant | Not supported. Use a single assistant | + +To run one: + +1. Follow the [Simulations quickstart](/observability/simulations-quickstart) and select your GPT-Live assistant as the target. +2. Write the scenario and success criteria. If the tester uses GPT-Live, choose one of its supported voices. +3. Run in **Voice** mode, then review the transcript, evaluations, and recording. + +Unmocked tools run for real. Use [tool mocks](/observability/simulations-advanced#mock-tool-responses), or point tools at a test service, before running scenarios that change data. A passing evaluation doesn't cover everything about the conversation, so listen to some recordings too. + +[Evals](/observability/evals-quickstart) check supported text-model decisions, such as which tool is called, using mock conversations. They don't test GPT-Live's listening, speech, or turn-taking. Use Voice Simulations to test the spoken conversation. + +## Watch production calls + +GPT-Live supports Vapi's post-call analysis and monitoring features: + +- **Extract data from each call.** Create [structured outputs](/assistants/structured-outputs-quickstart) and attach their IDs in `artifactPlan.structuredOutputIds`. Read the results in `call.artifact.structuredOutputs`. +- **Grade calls.** Create a [scorecard](/observability/scorecard-quickstart) from boolean or number structured outputs, and link it to the assistant. Read the grades in `call.artifact.scorecards`. +- **Receive call records on your server.** Set a Server URL and include `end-of-call-report` in the [server messages](/server-url/events). Your server receives an `end-of-call-report` message after the call. +- **Get alerts.** Create a [monitor](/observability/monitoring-quickstart) with its conditions, thresholds, evaluation windows, and notification destinations. Review its issues and alerts in Monitoring. + +A few things to know when setting these up: + +- When you update an assistant, keep its existing structured output IDs and artifact settings in the update, so you don't drop them. +- Existing [call analysis](/assistants/call-analysis) (`analysisPlan`) configurations keep working. For new setups, use structured outputs. +- A scorecard doesn't create a monitor. Set up monitors separately. +- [Authenticate your Server URL](/server-url/server-authentication). An end-of-call report can be delivered more than once, so handle repeats by call ID. +- [Boards](/observability/boards-quickstart) show call metrics and outcome trends for your assistant. + +A structured output is a good way to catch the problems on this page after the fact. For example, a boolean output: "Did the assistant tell the caller an action was completed that no tool result in the call confirmed?" A scorecard can then flag those calls. + +Some limits apply: + +- Transcript-based checks need transcripts, and audio-based checks need a recording. +- If an assistant has a recording-consent plan, recording is disabled, because GPT-Live doesn't collect that consent. +- Monitoring isn't available with Zero Data Retention. +- Live listening and monitor sockets aren't available. Post-call Monitoring doesn't provide live audio. +- Classic transcriber and text-to-speech timing metrics don't describe GPT-Live calls. + +## Latency + +Callers notice two kinds of delay, and they have different causes: + +- **Time until the caller hears something useful.** An acknowledgment, a relevant question, or an answer. This depends on the model's turn-taking, the call connection, and your speaker prompt, including whether it uses the wait well. +- **Time until the task is done correctly.** The result, stated accurately. This depends on when delegation happens, the reasoner, and your tools. + +Measure them separately. The call's messages and your service's logs have timestamps for delegations and tool calls, and the recording shows when the caller heard each part. + +When the task takes too long, try these, one at a time, and check that the outcome is still correct after each change: + +| Cause | What to try | +| --- | --- | +| Work starts late | Add a delegation trigger for the moment the inputs are known | +| Extra round trips | Return enough in one result to answer likely follow-ups, such as a whole day's times | +| Repeated lookups | Tell the reasoner when an earlier result still answers the request | +| Slow tools | Speed up your handler. Check the service log for the time each call takes | +| Reasoning time | Try a lower `model.reasoner.reasoningEffort`, such as `none`. The default is `low` | +| Model choice | Compare `gpt-5.6-luna`, `gpt-5.6-terra`, and `gpt-5.6-sol` on your scenarios | + +Lower effort or a smaller model can miss steps that a larger one handles. The right setting is the fastest one that still passes your scenarios. + +## Cost + +A GPT-Live call's cost has several parts: + +- **GPT-Live voice time**, billed per second for the length of the call, including silence and time spent waiting on the reasoner. +- **Vapi's platform fee**, per minute. +- **The reasoner**, billed by the tokens each delegation uses. Delegations include conversation context, so long conversations and frequent delegations can add input usage. +- **Telephony** and any other services, such as a phone number provider. + +Here's an illustrative estimate. It uses public rates as of September 30, 2026 and assumed token counts, so check current [Vapi pricing](https://vapi.ai/pricing) and your own calls before relying on it. + +| Part | Assumption | Cost | +| --- | --- | --- | +| GPT-Live voice | 4 minutes at $0.05 per minute | $0.20 | +| Vapi platform fee | 4 minutes at $0.05 per minute | $0.20 | +| Reasoner, `gpt-5.6-terra` | 8 delegations, about 4,000 uncached input and 150 output tokens each, at \$2 per million input and \$12 per million output tokens | about $0.08 | +| Telephony | Depends on your provider | not included | +| **Total** | | **about $0.48 plus telephony** | + +Your reasoner usage depends on the length of your prompts, tool results, and conversation, and on how often the assistant delegates. + +To see what a real call cost, retrieve it after it ends: + +```bash +curl --fail-with-body "https://api.vapi.ai/call/$CALL_ID" \ + -H "Authorization: Bearer $VAPI_PRIVATE_API_KEY" +``` + +The call's `costs` include the model cost with its billable voice time in `seconds`. If `usageComplete` is `false`, some usage wasn't reported and the cost may be understated. + +## Troubleshooting + +### GPT-Live isn't available, or calls don't start + +- **GPT-Live isn't offered for your organization.** It must be enabled for your organization. See [Access](/gpt-live/overview#access). +- **You use your own OpenAI API key.** The key needs access to GPT-Live and to the reasoner model you selected. +- **The assistant is rejected when you save it.** See [A voice or tool is rejected when saving](#a-voice-or-tool-is-rejected-when-saving). +- **Calls stopped starting after you added a saved tool.** Check that every saved tool ID the assistant references still exists. + +### The assistant said it would act, but nothing happened + +Follow the request through the call, in order: + +1. **Was there a tool call?** Check the call's messages at that point. If there's none, the speaker may not have delegated, or the reasoner may have returned a question instead of acting. Add a concrete trigger for that request to the speaker's delegation policy, and make answering the reasoner's questions a trigger too. +2. **Did the right tool run?** If a tool is missing, check that it's attached and that the reasoner prompt says when to use it. +3. **What did the tool return?** An error in the result means the action failed. The reasoner prompt should report failures plainly. +4. **What does your service show?** Its state is the answer to whether the action happened. + +Delegation is a model decision, so check over several calls after changing the prompt. + +### A tool result never reaches the conversation + +The tool ran, but the assistant didn't use its result. Check your webhook's response against [the tool response contract](/gpt-live/tools#the-tool-response-contract): HTTP 200, one entry per call with the `toolCallId` from the request, and `result` or `error` as a string. For an async tool, the result must come back in the response to the original request. A job that finishes after your webhook has responded needs a status tool. See [Slow and external work](/gpt-live/tools#slow-and-external-work). + +### Said goodbye but the call didn't end + +A spoken goodbye doesn't end the call. For an assistant-initiated hangup, check the call's messages for an `endCall` tool call. If there's none, list ending the call in the speaker's delegation triggers, tell the reasoner to call `endCall`, and make sure the tool is attached. If there is one, check its result. `maxDurationSeconds` and a final [idle hook](/gpt-live/design#when-the-caller-goes-quiet) with `endCall` are backstops for calls that stay open. + +### Answers are slow + +Work out which delay it is, using [Latency](#latency). For a long silence before any response, compare the recording with your speaker prompt and connection. For a long wait for a result, check the timestamps for when delegation started, when each tool call started and finished, and when the result was spoken. + +### A tool received the wrong value + +The reasoner works from the transcript. Compare the tool arguments with the recording. If the transcript misheard a name, number, or date, have the assistant [read back](/gpt-live/design#names-numbers-and-dates) important details before acting, and ask callers to spell names. + +### The assistant used an outdated result + +A correction arrived while earlier work was running. Check that the speaker prompt says to delegate the updated request and not to present old results, and that your results include the request details, such as the date, so a stale one is recognizable. + +### Waits feel awkward + +Listen to the recording. Repeated "still checking" messages, invented progress, or unrelated questions point to the speaker's guidance for [while work is running](/gpt-live/design#while-work-is-running). Long silences with no acknowledgment point the other way. Adjust one line at a time. + +### Check-ins happen at the wrong time + +Check-ins come from `customer.speech.timeout` hooks. If they fire while an async tool's result or an external job is still pending, that's expected: check-ins can resume once an async tool has been dispatched. Tell the caller about the wait and lengthen the timeouts. If check-ins come sooner than you expected after the caller asked for a moment, remember that the assistant's reply starts a new quiet period. If a check-in fires fewer times than you expect over a call, check `triggerMaxCount` and `triggerResetMode`, which control how many times a hook can fire. See [idle-message hooks](/gpt-live/configuration#idle-messages). + +### The assistant doesn't greet the caller + +Set `firstMessageMode` to `assistant-speaks-first` and give a text `firstMessage`. Audio greetings aren't supported, and without usable text the assistant waits for the caller. + +### A voice or tool is rejected when saving + +Use `voice.provider: "openai"` with a [supported voice](/gpt-live/configuration#voices), and remove voice and model fallbacks. Use supported tool types with unique names, and check saved tools too. Disable keypad input. + +### WebSocket calls fail to start or sound distorted + +Set the transport audio format explicitly to raw `pcm_s16le` at `24000` Hz, and send mono 16-bit little-endian audio at that rate. See [Call connections](/gpt-live/configuration#connect-a-call). + +### A transfer fails or doesn't connect + +Transfers work only on native Twilio and Vapi SIP calls, and only as blind transfers. A successful transfer request means the carrier accepted it, not that the destination answered. Check the destination and your carrier's requirements. + +### There's no recording or analysis result + +Check that recording is enabled and that no recording-consent plan is configured. A missing analysis result isn't a pass: check that the output IDs are attached and that the call has the transcript or recording the analysis needs. diff --git a/fern/gpt-live/tools.mdx b/fern/gpt-live/tools.mdx new file mode 100644 index 000000000..87e25df3f --- /dev/null +++ b/fern/gpt-live/tools.mdx @@ -0,0 +1,137 @@ +--- +title: Build with tools +subtitle: Connect your service so the assistant can look things up, take actions, and answer questions +description: How GPT-Live uses tools. Design tool inputs and results, handle actions that change state, return results correctly, handle slow and external work, and connect retrieval, transfers, and MCP tools. +slug: gpt-live/tools +--- + +In a GPT-Live assistant, the reasoner calls your tools. The speaker never calls them directly: it delegates, the reasoner decides which tool to use, and it returns what it found for the speaker to use in the conversation. Your service decides what actually happens. + +This page uses an appointment assistant as its example. It has four tools of its own, plus Vapi's built-in `endCall`: + +| Tool | What it does | Changes state? | +| --- | --- | --- | +| `lookupAvailability` | Lists every open time for a service, location, and date | No | +| `getServiceInfo` | Returns today's date, services, what to bring, locations, hours, and the change policy | No | +| `bookAppointment` | Books a time returned by a lookup | Yes | +| `cancelAppointment` | Cancels a booking | Yes | + +Function tools are covered in general in [Function tools](/tools/custom-tools). This page covers what matters when a GPT-Live assistant uses them. + +## Before you start + +- A GPT-Live assistant. The [Quickstart](/gpt-live/quickstart) creates one. +- For custom function tools, a server that handles Vapi's tool calls, protected with a [credential](/server-url/server-authentication). +- Speaker and reasoner prompts that say when to delegate and how to use each tool. [Design conversations](/gpt-live/design#turn-the-design-into-prompts) shows a full pair for this example. + +## Design tool inputs and results + +A tool's inputs and results shape the conversation as much as the prompts do. + +**Return enough to answer the likely follow-up.** The example lookup takes a service, location, and date, and returns every open time for that day. If the caller adds "ideally after three" while it runs, the result still answers the request, and the assistant can filter it without another lookup. A different day or location changes the request, so it needs a new lookup. Keep results short even when they're complete. + +**Echo the request details.** Include the date and location in each result. When the caller changes their mind during a lookup, the earlier result may still come back, and the details make it easy to recognize as out of date. + +**Say what didn't happen, as well as what did.** "Availability only. Nothing has been booked" helps the assistant avoid overstating a lookup. + +**Keep results plain.** The reasoner passes what matters to the speaker, so leave out Markdown and long explanations. + +**Make failures specific**, with the next step: "That time is no longer available. Look up availability again." + +## The tool response contract + +Vapi sends a `tool-calls` [server message](/server-url/events) to your server URL. Read `message.toolCallList`, run each call, and respond with HTTP 200 and one entry per call: + +```json +{ + "results": [ + { + "toolCallId": "call_1", + "result": "{\"status\":\"booked\",\"date\":\"2026-10-06\",\"time\":\"14:30\",\"location\":\"Downtown\"}" + } + ] +} +``` + +Use the `toolCallId` from the request. Use `result` for success and `error` for failure. Both must be flat strings, so serialize structured data. Return HTTP 200 even when a call fails, and report the failure in `error`. See [Function tools](/tools/custom-tools#server-response-format-providing-results-and-context) for the full contract. + +## Actions that change state + +Lookups are safe to repeat. Bookings and cancellations aren't, so they need more care from both the prompts and your service. + +### Get the caller's agreement first + +A caller choosing a time is a selection, not agreement to book. Have the speaker read back the details and ask, and have the reasoner book only after the caller clearly says yes. [Design conversations](/gpt-live/design#example-split-a-scheduling-prompt) shows the prompt wording. + +A tool call alone doesn't establish whether the caller heard the details and agreed to them. Prompts guide the model, but they don't enforce approval. If an action needs a stronger confirmation, enforce it in your service. + +### Check prerequisites in your service + +Before an action runs, check what your service can: + +- The time came from a real lookup and is still open. Someone else may have taken it. +- The request doesn't conflict with a booking already made, such as a second booking for the same caller when they meant to move the first. +- A repeated request doesn't create a second booking. Use the `toolCallId` or your own operation ID to recognize a retry, and return the original outcome. + +Tool calls can arrive together. The reasoner can request several at once, and they can run concurrently, so enforce any required order, such as lookup before booking, in the service. + +### Handle uncertain outcomes + +A request can time out after your service has already made the booking. Don't report "nothing was booked" from a timeout alone. Check the booking record before retrying, and reuse the same operation ID so a retry can't book twice. When the outcome is unknown, the assistant should say so and offer to check. + +### Changes to a completed action + +If your service moves a booking in two operations—cancel the old one, then book the new one—account for what happens if only the first succeeds. Have the reasoner look up the new time and get the caller's agreement first, then follow your service's change procedure. If the new booking fails after the cancellation, the assistant should say exactly that: the original booking was cancelled and the new time wasn't booked. + +Interrupting the assistant doesn't cancel work already running. See [When callers change the request](/gpt-live/design#when-callers-change-the-request). + +## Slow and external work + +A normal function tool makes the reasoner wait for your webhook's response. The conversation carries on meanwhile, as described in [While work is running](/gpt-live/design#while-work-is-running), and the reasoner returns the result for the speaker to use. This is the right choice when the caller is waiting for the answer. + +Set `async: true` on a function tool when the reasoner shouldn't wait for it, for example when the result doesn't affect what happens next in the call. Vapi records the call as pending and the reasoner carries on. Your webhook still returns the final result in its response to the original request, matched by `toolCallId`. When that response arrives, the result is added to the reasoner's history and given to the speaker as quiet context. The speaker can use it when it's relevant, but the result isn't guaranteed to be announced as soon as it arrives. If the caller needs to hear it, use a normal tool. + +Work that finishes after your webhook has responded is different. If your handler returns "queued" and a job keeps running elsewhere, nothing updates the conversation when the job finishes. Give the reasoner a status tool, such as `checkOrderStatus`, and tell it when to use it. Report a queued job as queued, never as done. + +Ending the call doesn't undo work your service has already started. Don't assume a hangup cancels an in-flight request; check its status in your service. + +While an async call or external job is pending, idle check-ins can resume if the caller goes quiet. See [When the caller goes quiet](/gpt-live/design#when-the-caller-goes-quiet). + +## Answer questions from your content + +`getServiceInfo` is a simple retrieval tool: the reasoner calls it when the caller asks about services, hours, or policies, and answers from the result. For a larger body of content, have the tool accept the caller's question and return only the most relevant passages, kept short. + +Knowledge bases and the Query tool aren't available to GPT-Live. Connect your content through a function, API request, or MCP tool instead. + +## Transfer to a person + +On native Twilio and Vapi SIP calls, the reasoner can transfer the caller with a `transferCall` tool. Give each destination a description so the reasoner knows when to choose it: + +```json +{ + "type": "transferCall", + "destinations": [ + { + "type": "number", + "number": "+14155550100", + "description": "Transfer to the front desk when the caller asks for a person or needs help the assistant can't give." + } + ] +} +``` + +Transfers are blind. Browser calls, including the dashboard's **Talk**, and raw WebSocket calls can't be transferred, so test this on a phone number. + +Add the transfer to the speaker's delegation triggers, and ask the speaker to tell the caller before it happens, for example "I'll put you through to the front desk now." That acknowledgment isn't guaranteed to finish before the transfer starts. + +A transfer involves separate events: the assistant says it's transferring, the reasoner requests the transfer, the carrier accepts it, and the destination answers. A successful request means the carrier accepted the transfer, not that someone answered. See [Settings and compatibility](/gpt-live/configuration#transfers-and-keypad-input) for supported options. + +## Use existing MCP tools + +MCP tools you already use with a Vapi assistant work with GPT-Live, with their existing server URLs, headers, and credentials. Attach the same saved tool IDs through `model.toolIds`, or keep inline `type: "mcp"` definitions in `model.tools`. See [MCP tools](/tools/mcp) for setup. + +At the start of a call, Vapi connects to the MCP server and makes its tools available to the reasoner. The tool list stays fixed for that call, so start a new call to pick up changes on the server. Put tool-use procedures in the reasoner prompt and delegation triggers in the speaker prompt. + +## Check the outcome, not just the conversation + +When you test a tool, look at three things together: what the assistant said, the tool calls and results in the call's messages, and your service's own records. The assistant saying a booking is done doesn't prove it happened. [Test and improve](/gpt-live/testing#scenarios-to-run) has a set of scenarios to run, including corrections, failures, and slow tools.