From 95126c4c81081f7923e3685cc4b0cb2fca947a35 Mon Sep 17 00:00:00 2001 From: Scott Lowe Date: Mon, 5 Oct 2026 20:14:38 -0700 Subject: [PATCH 1/2] Document latency limits for voice simulations Voice simulations can now fail when the assistant responds too slowly (VapiAI/vapi#21959, #21962, #21963; TEST-140). Add a "Set latency limits" section to Simulations advanced covering the Dashboard and API setup, the turn, model and voice metrics, the aggregations, when limits are skipped (chat mode, GPT-Live assistants) and how to read the results. Mention latency limits where the overview, manage and GPT-Live testing pages describe how a simulation is scored. Co-Authored-By: Claude Opus 5.5 --- fern/gpt-live/testing.mdx | 2 + fern/observability/simulations-advanced.mdx | 81 ++++++++++++++++++++- fern/observability/simulations-manage.mdx | 7 +- fern/observability/simulations-overview.mdx | 6 +- 4 files changed, 88 insertions(+), 8 deletions(-) diff --git a/fern/gpt-live/testing.mdx b/fern/gpt-live/testing.mdx index cbca438ce..6180ba4d4 100644 --- a/fern/gpt-live/testing.mdx +++ b/fern/gpt-live/testing.mdx @@ -62,6 +62,8 @@ To run one: Unmocked tools run for real. Use [tool mocks](/observability/simulations-advanced#mock-tool-responses), or point tools at a test service, before running scenarios that change data. A passing evaluation doesn't cover everything about the conversation, so listen to some recordings too. +[Latency limits](/observability/simulations-advanced#set-latency-limits) are skipped when the assistant under test uses GPT-Live, because its per-turn latency isn't measured yet. Measure GPT-Live latency with the call's timestamps, as described in [Latency](#latency). + [Evals](/observability/evals-quickstart) check supported text-model decisions, such as which tool is called, using mock conversations. They don't test GPT-Live's listening, speech, or turn-taking. Use Voice Simulations to test the spoken conversation. ## Watch production calls diff --git a/fern/observability/simulations-advanced.mdx b/fern/observability/simulations-advanced.mdx index eabad227f..3c7ef9c04 100644 --- a/fern/observability/simulations-advanced.mdx +++ b/fern/observability/simulations-advanced.mdx @@ -1,11 +1,11 @@ --- title: Simulations advanced -subtitle: Mock tools, send lifecycle webhooks, and reuse structured outputs in simulations -description: "Configure advanced simulations with variables, tool mocks, lifecycle webhooks, and reusable structured outputs for consistent testing." +subtitle: Mock tools, send lifecycle webhooks, set latency limits, and reuse structured outputs in simulations +description: "Configure advanced simulations with variables, tool mocks, lifecycle webhooks, latency limits, and reusable structured outputs for consistent testing." slug: observability/simulations-advanced --- -Advanced simulation options let you set variable values, mock tool responses, trigger lifecycle webhooks, and reuse structured outputs. Use them after you complete the [**Simulations quickstart**](/observability/simulations-quickstart) and need more control over behavior or test conditions. +Advanced simulation options let you set variable values, mock tool responses, trigger lifecycle webhooks, set latency limits, and reuse structured outputs. Use them after you complete the [**Simulations quickstart**](/observability/simulations-quickstart) and need more control over behavior or test conditions. ## How it works @@ -29,6 +29,9 @@ Open a suite and select **Edit**, then **Next**. The review step contains the ** Apply the same evaluation criterion across simulations. + + Fail a voice simulation when the assistant responds too slowly. + Build coverage with smoke tests, regression tests, and edge cases. @@ -239,6 +242,76 @@ curl -X PATCH "https://api.vapi.ai/eval/simulation/scenario/" \ +## Set latency limits + +A latency limit fails a voice simulation when the [**assistant**](/assistants) or [**squad**](/squads) under test responds too slowly. Vapi measures latency on each turn of the call, combines the turns into one number, and passes the limit when that number is at or below the limit you set. + +Latency limits work alongside structured-output evaluations. A scenario still needs at least one evaluation, and a required limit that fails marks the simulation as failed even when every evaluation passes. + + + + +On the **Success criteria** tab, under **Latency expectations**, select **Add latency expectation**. Configure each limit with four fields: + +- **Aggregation**: How the call's turns are combined: **Median**, **Mean**, **P95**, or **Max**. +- **Metric**: **Turn latency**, **Model latency**, or **Voice latency**. +- **Limit**: The maximum in milliseconds, as a whole number from 1 to 60,000. +- **Required**: Turn off to make the limit informational. An optional limit is reported but never fails the simulation. + +A new limit starts as a required median turn latency of 1,200 ms. Each simulation can have up to 20 limits. + + + + +Latency limits are the `latencyExpectations` property of the scenario. Add them when you create a scenario, or patch an existing one: + +```bash +curl -X PATCH "https://api.vapi.ai/eval/simulation/scenario/" \ + -H "Authorization: Bearer $VAPI_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "latencyExpectations": [ + { "metric": "turn", "aggregation": "median", "thresholdMs": 1200 }, + { "metric": "model", "aggregation": "p95", "thresholdMs": 800, "required": false } + ] + }' +``` + +`required` defaults to `true`. Sending `latencyExpectations` replaces the scenario's existing limits, and an empty array removes them. + + + + +### Metrics and aggregations + +| Metric | What it measures | +| --- | --- | +| `turn` | Total time from the end of the AI tester's speech to the start of the assistant's reply | +| `model` | Time for the model to return its first token | +| `voice` | Time for the voice provider to return the first audio | + +The aggregation reduces the call's turns to one value: `mean`, `median`, `p95`, or `max`. `p95` uses the nearest-rank method, so on a call with fewer than 20 turns it equals `max`. Measured values are rounded to the nearest millisecond before they are compared with the limit. + + + Gate on median turn latency. A short simulation has only a few turns, so `p95` and `max` can fail on a single slow turn. To watch the slowest turns without failing runs, add an optional `p95` or `max` limit next to a required `median` limit on the same metric. + + +### When a limit is skipped + +| Situation | Result | +| --- | --- | +| The run uses chat mode | Skipped. Chat runs have no audio, so there is no latency to measure. | +| The assistant under test uses [**GPT-Live**](/gpt-live/overview) | Skipped. Per-turn latency isn't measured for GPT-Live assistants yet. A GPT-Live AI tester talking to a classic assistant is still measured. | +| A voice run where no turn measured the metric | The limit fails. A latency that couldn't be measured never counts as within the limit. | + +Skipped limits don't affect the result. + +### Review latency results + +Open a run and select a simulation. The **Latency** section shows the average turn latency, then each limit with its measured value, its limit, and the number of turns measured. Optional and skipped limits are labeled. When a simulation fails, the count in the results list, such as `1 of 3 evaluations failed`, includes its latency limits. + +Through the API, each run item's `results.latencyEvaluations` holds one entry per limit with `actualMs`, `thresholdMs`, `sampleCount`, `passed`, `required`, and, for skipped limits, `isSkipped` and `skipReason`. `results.passed` accounts for required latency limits. + ## Set variable values for the assistant or squad Variables provide values for the [**dynamic variables**](/assistants/dynamic-variables) used by the [**assistant**](/assistants) or [**squad**](/squads) during the simulation. On the **Variables** tab, add **Name** and **Value** pairs. Each name must match a `{{variable}}` placeholder in the assistant's prompt, and the value is substituted during the run. @@ -255,6 +328,8 @@ Use variables to test specific inputs, such as a customer name or account tier, | There is no audio or recording | Check whether the run used chat mode. Use voice mode to test or record audio. | | Start or end webhooks do not trigger | Confirm the corresponding hook and its URL are configured on the scenario. | | A tool mock does not apply | Confirm that `toolName` matches the configured tool name exactly and that `enabled` is `true`. | +| A latency limit fails without a measured value | No turn in the call measured that metric. Confirm the run used voice mode and that the conversation completed at least one exchange. | +| Latency limits show as skipped | Check whether the run used chat mode or the assistant under test uses GPT-Live. Both skip latency limits. | ## Next steps diff --git a/fern/observability/simulations-manage.mdx b/fern/observability/simulations-manage.mdx index 7169e6802..9bfb4c14f 100644 --- a/fern/observability/simulations-manage.mdx +++ b/fern/observability/simulations-manage.mdx @@ -73,7 +73,7 @@ curl -X PATCH "https://api.vapi.ai/eval/simulation/personality/" Open **Simulations**, then select **Runs** to see every run with its suite, iterations, [**assistant**](/assistants) or [**squad**](/squads), run date, and overall result (`Passed` or a failed count such as `1/1 failed`). Filter by **time range**, **status**, or **assistant or squad** to find a run. -Open a run to see whether each simulation passed or failed, its evaluations, and its transcript. Open an active run to watch the transcript and evaluations update. If a run or iteration can't complete, it shows a **Failure reason** describing what went wrong, such as no payment method on file. See the [**Simulations quickstart**](/observability/simulations-quickstart) for the full walkthrough. +Open a run to see whether each simulation passed or failed, its evaluations, its latency results for voice runs, and its transcript. Open an active run to watch the transcript and evaluations update. If a run or iteration can't complete, it shows a **Failure reason** describing what went wrong, such as no payment method on file. See the [**Simulations quickstart**](/observability/simulations-quickstart) for the full walkthrough. @@ -90,7 +90,7 @@ curl -X GET "https://api.vapi.ai/eval/simulation/run/" \ -H "Authorization: Bearer $VAPI_API_KEY" ``` -The run includes `status` and `itemCounts` (`total`, `passed`, `failed`, and so on). Fetch the result for each simulation iteration with `GET /eval/simulation/run//item`. +The run includes `status` and `itemCounts` (`total`, `passed`, `failed`, and so on). Fetch the result for each simulation iteration with `GET /eval/simulation/run//item`. Each item's `results` holds its `evaluations` and, when the scenario sets latency limits, its `latencyEvaluations`. See [**Review latency results**](/observability/simulations-advanced#review-latency-results). @@ -144,12 +144,15 @@ Test realistic variations such as an ambiguous request, an impatient customer, a Keep each evaluation focused on one observable outcome. Use a descriptive name, choose a Boolean or numeric value that can be measured consistently, and avoid combining several independent requirements into one evaluation. +To catch slow responses, add a [**latency limit**](/observability/simulations-advanced#set-latency-limits) rather than an evaluation, and run the suite in voice mode. Median turn latency is the most stable limit to gate a release on. + ### Choose voice or chat mode | Testing goal | Mode | Why | | --- | --- | --- | | Iterate on prompts, tools, and conversation logic | Chat (`vapi.webchat`) | Runs without audio processing, so it is faster and costs less. | | Investigate speech recognition, voice output, or interruptions | Voice (`vapi.websocket`) | Exercises synthetic audio. Review the recording, not just the pass/fail label. | +| Gate on response latency | Voice (`vapi.websocket`) | Latency limits are skipped in chat mode. | | Receive call-specific webhook data | Voice (`vapi.websocket`) | Start and end webhooks fire in both modes, but chat payloads omit call-specific fields. | | Check representative voice journeys before launch | Voice (`vapi.websocket`) | Adds voice coverage. Also make controlled calls through the real phone path and sandbox integrations. | diff --git a/fern/observability/simulations-overview.mdx b/fern/observability/simulations-overview.mdx index d5c04ada0..96d70eabd 100644 --- a/fern/observability/simulations-overview.mdx +++ b/fern/observability/simulations-overview.mdx @@ -13,7 +13,7 @@ Simulations are automated tests that run an AI tester through a real conversatio | -- | -- | | **Simulation suite** | A simulation suite groups one or more simulations that you can run against one or more [**assistants**](/assistants) or [**squads**](/squads). | | **Simulation** | A simulation pairs a scenario with a personality to test one situation. | -| **Scenario** | A scenario defines the AI tester's intent and the success criteria that determine the result. | +| **Scenario** | A scenario defines the AI tester's intent and the success criteria that determine the result: structured-output evaluations and, for voice runs, any latency limits. | | **Personality** | A personality defines how the AI tester behaves, including its model, transcriber, and voice. | | **AI tester** | An AI tester drives the simulated conversation according to a scenario and personality. | @@ -54,7 +54,7 @@ directly instead of synthesized speech and transcription. Use **chat mode** for rapid iteration during development, then switch to **voice mode** for final validation. -When you run the suite, Vapi runs each simulation for the number of iterations you select. During each iteration, the AI tester holds a live conversation with the assistant or squad. Vapi then evaluates each conversation with [**structured outputs**](/assistants/structured-outputs-quickstart) and reports which criteria passed or failed. +When you run the suite, Vapi runs each simulation for the number of iterations you select. During each iteration, the AI tester holds a live conversation with the assistant or squad. Vapi then evaluates each conversation with [**structured outputs**](/assistants/structured-outputs-quickstart), checks any [**latency limits**](/observability/simulations-advanced#set-latency-limits) on voice runs, and reports which criteria passed or failed. ## When to use Simulations @@ -82,7 +82,7 @@ Simulations test live interactions. An AI tester with a defined personality and | **What you provide** | A scripted conversation with expected responses | An AI tester personality and an intent | | **How it runs** | Fixed-context, turn-by-turn checks | Dynamic; the AI tester improvises the conversation | | **Transport** | Chat (mock conversations) | Voice, or text-only | -| **Evaluation** | Exact match, regex, AI judge, tool-call checks | Structured outputs on the conversation outcome | +| **Evaluation** | Exact match, regex, AI judge, tool-call checks | Structured outputs on the conversation outcome, plus latency limits on voice runs | | **Reach for it when** | You need focused checks at a known point | You need realistic end-to-end or voice behavior | Choose Evals when you want to lock down a specific response or verify a tool call's arguments with fast, rerunnable checks. From d95a911c36e15125276f9ecbf988fdb0c6045282 Mon Sep 17 00:00:00 2001 From: scott-lowe-vapi Date: Tue, 6 Oct 2026 16:01:34 -0700 Subject: [PATCH 2/2] Update simulations-advanced.mdx --- fern/observability/simulations-advanced.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/fern/observability/simulations-advanced.mdx b/fern/observability/simulations-advanced.mdx index 3c7ef9c04..ccde3df4f 100644 --- a/fern/observability/simulations-advanced.mdx +++ b/fern/observability/simulations-advanced.mdx @@ -308,7 +308,7 @@ Skipped limits don't affect the result. ### Review latency results -Open a run and select a simulation. The **Latency** section shows the average turn latency, then each limit with its measured value, its limit, and the number of turns measured. Optional and skipped limits are labeled. When a simulation fails, the count in the results list, such as `1 of 3 evaluations failed`, includes its latency limits. +Open a run and select a simulation. The **Latency** section shows average turn latency when at least one turn was measured. Each limit shows its threshold and result. When at least one turn measured that limit's metric, the result also shows the measured value and number of turns. Optional and skipped limits are labeled. When a simulation fails, the count in the results list, such as `1 of 3 evaluations failed`, includes its latency limits. Through the API, each run item's `results.latencyEvaluations` holds one entry per limit with `actualMs`, `thresholdMs`, `sampleCount`, `passed`, `required`, and, for skipped limits, `isSkipped` and `skipReason`. `results.passed` accounts for required latency limits.