Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions fern/gpt-live/testing.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,8 @@ To run one:

Unmocked tools run for real. Use [tool mocks](/observability/simulations-advanced#mock-tool-responses), or point tools at a test service, before running scenarios that change data. A passing evaluation doesn't cover everything about the conversation, so listen to some recordings too.

[Latency limits](/observability/simulations-advanced#set-latency-limits) are skipped when the assistant under test uses GPT-Live, because its per-turn latency isn't measured yet. Measure GPT-Live latency with the call's timestamps, as described in [Latency](#latency).

[Evals](/observability/evals-quickstart) check supported text-model decisions, such as which tool is called, using mock conversations. They don't test GPT-Live's listening, speech, or turn-taking. Use Voice Simulations to test the spoken conversation.

## Watch production calls
Expand Down
81 changes: 78 additions & 3 deletions fern/observability/simulations-advanced.mdx
Original file line number Diff line number Diff line change
@@ -1,11 +1,11 @@
---
title: Simulations advanced
subtitle: Mock tools, send lifecycle webhooks, and reuse structured outputs in simulations
description: "Configure advanced simulations with variables, tool mocks, lifecycle webhooks, and reusable structured outputs for consistent testing."
subtitle: Mock tools, send lifecycle webhooks, set latency limits, and reuse structured outputs in simulations
description: "Configure advanced simulations with variables, tool mocks, lifecycle webhooks, latency limits, and reusable structured outputs for consistent testing."
slug: observability/simulations-advanced
---

Advanced simulation options let you set variable values, mock tool responses, trigger lifecycle webhooks, and reuse structured outputs. Use them after you complete the [**Simulations quickstart**](/observability/simulations-quickstart) and need more control over behavior or test conditions.
Advanced simulation options let you set variable values, mock tool responses, trigger lifecycle webhooks, set latency limits, and reuse structured outputs. Use them after you complete the [**Simulations quickstart**](/observability/simulations-quickstart) and need more control over behavior or test conditions.

## How it works

Expand All @@ -29,6 +29,9 @@ Open a suite and select **Edit**, then **Next**. The review step contains the **
<Card title="Reuse structured outputs" icon="database" href="#reuse-structured-outputs">
Apply the same evaluation criterion across simulations.
</Card>
<Card title="Set latency limits" icon="stopwatch" href="#set-latency-limits">
Fail a voice simulation when the assistant responds too slowly.
</Card>
<Card title="Manage test coverage" icon="flask" href="/observability/simulations-manage#testing-strategies">
Build coverage with smoke tests, regression tests, and edge cases.
</Card>
Expand Down Expand Up @@ -239,6 +242,76 @@ curl -X PATCH "https://api.vapi.ai/eval/simulation/scenario/<scenario-id>" \
</Tab>
</Tabs>

## Set latency limits

A latency limit fails a voice simulation when the [**assistant**](/assistants) or [**squad**](/squads) under test responds too slowly. Vapi measures latency on each turn of the call, combines the turns into one number, and passes the limit when that number is at or below the limit you set.

Latency limits work alongside structured-output evaluations. A scenario still needs at least one evaluation, and a required limit that fails marks the simulation as failed even when every evaluation passes.

<Tabs>
<Tab title="Dashboard">

On the **Success criteria** tab, under **Latency expectations**, select **Add latency expectation**. Configure each limit with four fields:

- **Aggregation**: How the call's turns are combined: **Median**, **Mean**, **P95**, or **Max**.
- **Metric**: **Turn latency**, **Model latency**, or **Voice latency**.
- **Limit**: The maximum in milliseconds, as a whole number from 1 to 60,000.
- **Required**: Turn off to make the limit informational. An optional limit is reported but never fails the simulation.

A new limit starts as a required median turn latency of 1,200 ms. Each simulation can have up to 20 limits.

</Tab>
<Tab title="cURL">

Latency limits are the `latencyExpectations` property of the scenario. Add them when you create a scenario, or patch an existing one:

```bash
curl -X PATCH "https://api.vapi.ai/eval/simulation/scenario/<scenario-id>" \
-H "Authorization: Bearer $VAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"latencyExpectations": [
{ "metric": "turn", "aggregation": "median", "thresholdMs": 1200 },
{ "metric": "model", "aggregation": "p95", "thresholdMs": 800, "required": false }
]
}'
```

`required` defaults to `true`. Sending `latencyExpectations` replaces the scenario's existing limits, and an empty array removes them.

</Tab>
</Tabs>

### Metrics and aggregations

| Metric | What it measures |
| --- | --- |
| `turn` | Total time from the end of the AI tester's speech to the start of the assistant's reply |
| `model` | Time for the model to return its first token |
| `voice` | Time for the voice provider to return the first audio |

The aggregation reduces the call's turns to one value: `mean`, `median`, `p95`, or `max`. `p95` uses the nearest-rank method, so on a call with fewer than 20 turns it equals `max`. Measured values are rounded to the nearest millisecond before they are compared with the limit.

<Tip>
Gate on median turn latency. A short simulation has only a few turns, so `p95` and `max` can fail on a single slow turn. To watch the slowest turns without failing runs, add an optional `p95` or `max` limit next to a required `median` limit on the same metric.
</Tip>

### When a limit is skipped
Comment thread
stephenvapiai marked this conversation as resolved.

| Situation | Result |
| --- | --- |
| The run uses chat mode | Skipped. Chat runs have no audio, so there is no latency to measure. |
| The assistant under test uses [**GPT-Live**](/gpt-live/overview) | Skipped. Per-turn latency isn't measured for GPT-Live assistants yet. A GPT-Live AI tester talking to a classic assistant is still measured. |
| A voice run where no turn measured the metric | The limit fails. A latency that couldn't be measured never counts as within the limit. |

Skipped limits don't affect the result.

### Review latency results

Open a run and select a simulation. The **Latency** section shows average turn latency when at least one turn was measured. Each limit shows its threshold and result. When at least one turn measured that limit's metric, the result also shows the measured value and number of turns. Optional and skipped limits are labeled. When a simulation fails, the count in the results list, such as `1 of 3 evaluations failed`, includes its latency limits.

Through the API, each run item's `results.latencyEvaluations` holds one entry per limit with `actualMs`, `thresholdMs`, `sampleCount`, `passed`, `required`, and, for skipped limits, `isSkipped` and `skipReason`. `results.passed` accounts for required latency limits.

## Set variable values for the assistant or squad

Variables provide values for the [**dynamic variables**](/assistants/dynamic-variables) used by the [**assistant**](/assistants) or [**squad**](/squads) during the simulation. On the **Variables** tab, add **Name** and **Value** pairs. Each name must match a `{{variable}}` placeholder in the assistant's prompt, and the value is substituted during the run.
Expand All @@ -255,6 +328,8 @@ Use variables to test specific inputs, such as a customer name or account tier,
| There is no audio or recording | Check whether the run used chat mode. Use voice mode to test or record audio. |
| Start or end webhooks do not trigger | Confirm the corresponding hook and its URL are configured on the scenario. |
| A tool mock does not apply | Confirm that `toolName` matches the configured tool name exactly and that `enabled` is `true`. |
| A latency limit fails without a measured value | No turn in the call measured that metric. Confirm the run used voice mode and that the conversation completed at least one exchange. |
| Latency limits show as skipped | Check whether the run used chat mode or the assistant under test uses GPT-Live. Both skip latency limits. |

## Next steps

Expand Down
7 changes: 5 additions & 2 deletions fern/observability/simulations-manage.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -73,7 +73,7 @@ curl -X PATCH "https://api.vapi.ai/eval/simulation/personality/<personality-id>"

Open **Simulations**, then select **Runs** to see every run with its suite, iterations, [**assistant**](/assistants) or [**squad**](/squads), run date, and overall result (`Passed` or a failed count such as `1/1 failed`). Filter by **time range**, **status**, or **assistant or squad** to find a run.

Open a run to see whether each simulation passed or failed, its evaluations, and its transcript. Open an active run to watch the transcript and evaluations update. If a run or iteration can't complete, it shows a **Failure reason** describing what went wrong, such as no payment method on file. See the [**Simulations quickstart**](/observability/simulations-quickstart) for the full walkthrough.
Open a run to see whether each simulation passed or failed, its evaluations, its latency results for voice runs, and its transcript. Open an active run to watch the transcript and evaluations update. If a run or iteration can't complete, it shows a **Failure reason** describing what went wrong, such as no payment method on file. See the [**Simulations quickstart**](/observability/simulations-quickstart) for the full walkthrough.

</Tab>
<Tab title="cURL">
Expand All @@ -90,7 +90,7 @@ curl -X GET "https://api.vapi.ai/eval/simulation/run/<run-id>" \
-H "Authorization: Bearer $VAPI_API_KEY"
```

The run includes `status` and `itemCounts` (`total`, `passed`, `failed`, and so on). Fetch the result for each simulation iteration with `GET /eval/simulation/run/<run-id>/item`.
The run includes `status` and `itemCounts` (`total`, `passed`, `failed`, and so on). Fetch the result for each simulation iteration with `GET /eval/simulation/run/<run-id>/item`. Each item's `results` holds its `evaluations` and, when the scenario sets latency limits, its `latencyEvaluations`. See [**Review latency results**](/observability/simulations-advanced#review-latency-results).

</Tab>
</Tabs>
Expand Down Expand Up @@ -144,12 +144,15 @@ Test realistic variations such as an ambiguous request, an impatient customer, a

Keep each evaluation focused on one observable outcome. Use a descriptive name, choose a Boolean or numeric value that can be measured consistently, and avoid combining several independent requirements into one evaluation.

To catch slow responses, add a [**latency limit**](/observability/simulations-advanced#set-latency-limits) rather than an evaluation, and run the suite in voice mode. Median turn latency is the most stable limit to gate a release on.

### Choose voice or chat mode

| Testing goal | Mode | Why |
| --- | --- | --- |
| Iterate on prompts, tools, and conversation logic | Chat (`vapi.webchat`) | Runs without audio processing, so it is faster and costs less. |
| Investigate speech recognition, voice output, or interruptions | Voice (`vapi.websocket`) | Exercises synthetic audio. Review the recording, not just the pass/fail label. |
| Gate on response latency | Voice (`vapi.websocket`) | Latency limits are skipped in chat mode. |
| Receive call-specific webhook data | Voice (`vapi.websocket`) | Start and end webhooks fire in both modes, but chat payloads omit call-specific fields. |
| Check representative voice journeys before launch | Voice (`vapi.websocket`) | Adds voice coverage. Also make controlled calls through the real phone path and sandbox integrations. |

Expand Down
6 changes: 3 additions & 3 deletions fern/observability/simulations-overview.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ Simulations are automated tests that run an AI tester through a real conversatio
| -- | -- |
| **Simulation suite** | A simulation suite groups one or more simulations that you can run against one or more [**assistants**](/assistants) or [**squads**](/squads). |
| **Simulation** | A simulation pairs a scenario with a personality to test one situation. |
| **Scenario** | A scenario defines the AI tester's intent and the success criteria that determine the result. |
| **Scenario** | A scenario defines the AI tester's intent and the success criteria that determine the result: structured-output evaluations and, for voice runs, any latency limits. |
| **Personality** | A personality defines how the AI tester behaves, including its model, transcriber, and voice. |
| **AI tester** | An AI tester drives the simulated conversation according to a scenario and personality. |

Expand Down Expand Up @@ -54,7 +54,7 @@ directly instead of synthesized speech and transcription.
Use **chat mode** for rapid iteration during development, then switch to **voice mode** for final validation.
</Tip>

When you run the suite, Vapi runs each simulation for the number of iterations you select. During each iteration, the AI tester holds a live conversation with the assistant or squad. Vapi then evaluates each conversation with [**structured outputs**](/assistants/structured-outputs-quickstart) and reports which criteria passed or failed.
When you run the suite, Vapi runs each simulation for the number of iterations you select. During each iteration, the AI tester holds a live conversation with the assistant or squad. Vapi then evaluates each conversation with [**structured outputs**](/assistants/structured-outputs-quickstart), checks any [**latency limits**](/observability/simulations-advanced#set-latency-limits) on voice runs, and reports which criteria passed or failed.

## When to use Simulations

Expand Down Expand Up @@ -82,7 +82,7 @@ Simulations test live interactions. An AI tester with a defined personality and
| **What you provide** | A scripted conversation with expected responses | An AI tester personality and an intent |
| **How it runs** | Fixed-context, turn-by-turn checks | Dynamic; the AI tester improvises the conversation |
| **Transport** | Chat (mock conversations) | Voice, or text-only |
| **Evaluation** | Exact match, regex, AI judge, tool-call checks | Structured outputs on the conversation outcome |
| **Evaluation** | Exact match, regex, AI judge, tool-call checks | Structured outputs on the conversation outcome, plus latency limits on voice runs |
| **Reach for it when** | You need focused checks at a known point | You need realistic end-to-end or voice behavior |

Choose Evals when you want to lock down a specific response or verify a tool call's arguments with fast, rerunnable checks.
Expand Down
Loading