From 93df979c7f83c2e5f4cfc3c09df2f05b20fd2cdc Mon Sep 17 00:00:00 2001
From: "mintlify[bot]" <109931778+mintlify[bot]@users.noreply.github.com>
Date: Wed, 26 Aug 2026 20:25:29 +0000
Subject: [PATCH 1/3] docs: fix style and tone consistency in checkpointing
page
---
performance/checkpointing.mdx | 16 ++++++++--------
1 file changed, 8 insertions(+), 8 deletions(-)
diff --git a/performance/checkpointing.mdx b/performance/checkpointing.mdx
index ec0193ff..0d73b796 100644
--- a/performance/checkpointing.mdx
+++ b/performance/checkpointing.mdx
@@ -18,7 +18,7 @@ This is useful for both CPU-only and GPU workloads. For CPU applications, checkp
For example, ML and LLM frameworks often load large model weights and compile CUDA kernels at container start time, which can take many seconds or minutes. Loading from a checkpoint that already contains this initialized state can skip most of that delay.
-Since this feature is still in beta, please report all issues to the team via our [Discord Community](https://discord.gg/ATj6USmeE2) or via [Email](mailto:support@cerebrium.ai).
+Since this feature is still in beta, please report all issues to the team via our [Discord Community](https://discord.gg/ATj6USmeE2) or via [email](mailto:support@cerebrium.ai).
## How to use
@@ -33,7 +33,7 @@ checkpointing = true
### 2. Send the checkpoint trigger
-You have full control over when to create a checkpoint - preferably it's after the majority of your application's initialization work is complete and your desired state is reached. The state of your application at this exact moment is what will be restored for future container launches.
+You have full control over when to create a checkpoint. Trigger it after your application completes the majority of its initialization work and reaches your desired state. Cerebrium restores the state of your application at this exact moment for future container launches.
Once you reach this optimal point in your startup logic, send a POST request from inside the container to instruct Cerebrium to capture the checkpoint:
@@ -52,18 +52,18 @@ with urllib.request.urlopen(req, timeout=300) as response:
The endpoint uses `169.254.169.253`, a link-local address that routes to the Cerebrium runtime sidecar inside the container. The address is reachable only from inside the container, not from external networks.
-Set the HTTP client timeout to at least 300 seconds. Checkpoint duration scales with the amount of memory captured - from a few seconds for small CPU workloads to several minutes for large GPU snapshots.
+Set the HTTP client timeout to at least 300 seconds. Checkpoint duration scales with the amount of memory captured. It ranges from a few seconds for small CPU workloads to several minutes for large GPU snapshots.
### 3. When the runtime creates a checkpoint
-When the runtime receives the POST, it checks whether a new checkpoint is required. To save resources, the system skips checkpoint creation if:
+When the runtime receives the POST, it checks whether a new checkpoint is required. To save resources, the runtime skips checkpoint creation if:
1. A checkpoint already exists for the current build version.
2. Another container instance is already undergoing the checkpointing process.
-If a checkpoint should occur, the container is frozen for the duration of the process. GPU memory is copied to CPU memory, and then all container memory is written to storage. The saved checkpoint is then distributed throughout the region.
+If a checkpoint should occur, the runtime freezes the container for the duration of the process. It copies GPU memory to CPU memory, then writes all container memory to storage. Cerebrium then distributes the saved checkpoint throughout the region.
-Checkpoint size roughly equals CPU memory in use at trigger time plus GPU memory copied during the freeze. For GPU workloads, expect the snapshot to be on the order of model weights plus runtime overhead unless caches are dropped first. Allocate enough container memory to hold the GPU dump in addition to normal usage — see **Memory overhead** under Limitations.
+Checkpoint size roughly equals CPU memory in use at trigger time plus GPU memory copied during the freeze. For GPU workloads, expect the snapshot to be on the order of model weights plus runtime overhead unless caches are dropped first. Allocate enough container memory to hold the GPU dump in addition to normal usage. See **Memory overhead** under Limitations.
### 4. Verify restoration
@@ -73,7 +73,7 @@ If checkpoint creation succeeds, subsequent containers restore from that snapsho
A checkpoint is tightly coupled to a single deployment. To stop restoring from checkpoints, remove the POST request and redeploy the application.
-You can find several implementations in our [Examples repository on Github](https://github.com/CerebriumAI/examples).
+You can find several implementations in our [Examples repository on GitHub](https://github.com/CerebriumAI/examples).
### vLLM example
@@ -124,4 +124,4 @@ engine.wake_up()
vLLM checkpointing support is not complete but is still possible. See [vllm-project/vllm#34303](https://github.com/vllm-project/vllm/issues/34303) and related issues.
-The larger the size of the memory checkpoint, the slower the restore is. Reduce the size of the snapshot substantially and improve startup times by dropping the KV Cache before checkpoint and recreating it after restore. vLLM has functionality that does this built in as part of [vLLM Sleep Mode](https://docs.vllm.ai/en/latest/features/sleep_mode/).
+The larger the size of the memory checkpoint, the slower the restore is. Reduce the size of the snapshot substantially and improve startup times by dropping the KV cache before checkpoint and recreating it after restore. vLLM has functionality that does this built in as part of [vLLM Sleep Mode](https://docs.vllm.ai/en/latest/features/sleep_mode/).
From 01fd202f531471c16f1ed9fe7ebbaf9bb53354d1 Mon Sep 17 00:00:00 2001
From: "mintlify[bot]" <109931778+mintlify[bot]@users.noreply.github.com>
Date: Wed, 26 Aug 2026 20:32:14 +0000
Subject: [PATCH 2/3] docs: fix style and tone consistency across 28 pages from
SEO audit scope
---
container-images/custom-dockerfiles.mdx | 4 ++--
container-images/custom-web-servers.mdx | 2 +-
.../defining-container-images.mdx | 24 +++++++++----------
deployments/ci-cd.mdx | 4 ++--
deployments/multi-region-deployment.mdx | 8 +++----
endpoints/async.mdx | 8 +++----
endpoints/inference-api.mdx | 4 ++--
endpoints/streaming.mdx | 2 +-
endpoints/webhook.mdx | 6 ++---
endpoints/websockets.mdx | 4 ++--
hardware/cpu-and-memory.mdx | 6 ++---
hardware/using-cuda.mdx | 2 +-
hardware/using-gpus.mdx | 4 ++--
networking/custom-domains.mdx | 10 ++++----
other-topics/request-response-logging.mdx | 4 ++--
other-topics/using-secrets.mdx | 4 ++--
partner-services/deepgram.mdx | 8 +++----
partner-services/index.mdx | 4 ++--
partner-services/rime.mdx | 2 +-
scaling/batching-concurrency.mdx | 2 +-
scaling/graceful-termination.mdx | 2 +-
toml-reference/toml-reference.mdx | 4 ++--
v4/examples/comfyUI.mdx | 6 ++---
...oy-a-vision-language-model-with-sglang.mdx | 4 ++--
...y-an-llm-with-tensorrtllm-tritonserver.mdx | 4 ++--
v4/examples/gpt-oss.mdx | 4 ++--
v4/examples/langchain-langsmith.mdx | 10 ++++----
v4/examples/livekit-outbound-agent.mdx | 14 +++++------
.../openai-compatible-endpoint-vllm.mdx | 2 +-
v4/examples/realtime-voice-agents.mdx | 10 ++++----
v4/examples/sdxl.mdx | 2 +-
v4/examples/transcribe-whisper.mdx | 12 +++++-----
v4/examples/twilio-voice-agent.mdx | 6 ++---
v4/examples/wandb-sweep.mdx | 6 ++---
34 files changed, 99 insertions(+), 99 deletions(-)
diff --git a/container-images/custom-dockerfiles.mdx b/container-images/custom-dockerfiles.mdx
index 391ef113..5f898219 100644
--- a/container-images/custom-dockerfiles.mdx
+++ b/container-images/custom-dockerfiles.mdx
@@ -48,7 +48,7 @@ CMD ["python", "-m", "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8192
Dockerfiles for Cerebrium have three requirements:
-1. Expose a port with the `EXPOSE` command - this port is referenced in `cerebrium.toml`
+1. Expose a port with the `EXPOSE` command. This port is referenced in `cerebrium.toml`
2. Include a `CMD` command to specify the container's startup process (typically the server)
3. Set the working directory with `WORKDIR` to ensure correct file paths (defaults to root if not specified)
@@ -87,7 +87,7 @@ entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8192"]
## Building Generic Dockerized Apps
-Cerebrium supports non-Python apps as long as a Dockerfile is provided. The following example shows a Rust-based API server using the Axum framework:
+Cerebrium supports non-Python apps as long as you provide a Dockerfile. The following example shows a Rust-based API server using the Axum framework:
```rust
use axum::{
diff --git a/container-images/custom-web-servers.mdx b/container-images/custom-web-servers.mdx
index 6c210bb1..3949ba04 100644
--- a/container-images/custom-web-servers.mdx
+++ b/container-images/custom-web-servers.mdx
@@ -3,7 +3,7 @@ title: "Custom Python Web Servers"
description: Run FastAPI and other ASGI or WSGI Python web servers on Cerebrium with a custom runtime by setting the entrypoint, port, and health check endpoints.
---
-Cerebrium's default runtime covers most app needs. For more control, use ASGI or WSGI servers through the custom runtime feature - enabling custom authentication, dynamic batching, frontend dashboards, public endpoints, and WebSocket connections.
+Cerebrium's default runtime covers most app needs. For more control, use ASGI or WSGI servers through the custom runtime feature. This enables custom authentication, dynamic batching, frontend dashboards, public endpoints, and WebSocket connections.
## Setting Up Custom Servers
diff --git a/container-images/defining-container-images.mdx b/container-images/defining-container-images.mdx
index 2654d9cb..e0e43b64 100644
--- a/container-images/defining-container-images.mdx
+++ b/container-images/defining-container-images.mdx
@@ -5,7 +5,7 @@ description: Define your Cerebrium container image in cerebrium.toml, from Pytho
## Introduction
-Cerebrium abstracts infrastructure management into configuration, so teams focus on app code. A single TOML file manages environment setup, deployments, and scaling — tasks that typically require dedicated teams.
+Cerebrium abstracts infrastructure management into configuration, so teams focus on app code. A single TOML file manages environment setup, deployments, and scaling: tasks that typically require dedicated teams.
Unlike traditional Docker or Kubernetes setups with multiple configuration files and orchestration rules, Cerebrium uses a single `cerebrium.toml` file. The system handles container lifecycle, networking, and scaling automatically based on this configuration.
@@ -18,11 +18,11 @@ Python decorators scatter infrastructure settings throughout code files, making
Run `cerebrium init` to create a `cerebrium.toml` file in the project root. Edit it to match the app's requirements.
- It is possible to initialize an existing project by adding a `cerebrium.toml`
- file to the root of your codebase, defining your entrypoint (`main.py` if
- using the default runtime, or adding an entrypoint to the .toml file if using
- a custom runtime) and including the necessary files in the `deployment`
- section of your `cerebrium.toml` file.
+ You can initialize an existing project by adding a `cerebrium.toml` file to
+ the root of your codebase. Define your entrypoint (`main.py` if using the
+ default runtime, or add an entrypoint to the .toml file if using a custom
+ runtime) and include the necessary files in the `deployment` section of your
+ `cerebrium.toml` file.
## Hardware Configuration
@@ -39,7 +39,7 @@ gpu_count = 1 # Number of GPUs
For detailed hardware specifications see the [toml reference](/toml-reference/toml-reference#hardware-configuration).
-## Dependency management
+## Dependency Management
### Selecting a Python Version
@@ -86,7 +86,7 @@ Cerebrium caches pip packages at the node level - including wheel files and comp
### Adding APT Packages
-System-level packages (image-processing libraries, audio codecs, etc.) are declared under `[cerebrium.dependencies.apt]`:
+Declare system-level packages (image-processing libraries, audio codecs, etc.) under `[cerebrium.dependencies.apt]`:
```toml
[cerebrium.dependencies.apt]
@@ -158,7 +158,7 @@ shell_commands = [
]
```
-Use shell commands for tasks that require the fully configured environment — such as compiling code that depends on installed libraries or downloading resources.
+Use shell commands for tasks that require the fully configured environment, such as compiling code that depends on installed libraries or downloading resources.
## Custom Docker Base Images
@@ -271,14 +271,14 @@ vllm = "latest"
### Important Notes
-- Code is mounted in `/cortex` - adjust paths accordingly.
+- Code is mounted in `/cortex`. Adjust paths accordingly.
- The port in your entrypoint must match the `port` parameter.
- Install any required server packages (uvicorn, gunicorn, etc.) via pip dependencies.
- All endpoints will be available at `https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/your/endpoint`.
-Deploy with `cerebrium deploy -y` - the system automatically detects custom runtime configuration.
+Deploy with `cerebrium deploy -y`. The system automatically detects custom runtime configuration.
-## Deployment process
+## Deployment Process

diff --git a/deployments/ci-cd.mdx b/deployments/ci-cd.mdx
index 12b1de45..c48effdc 100644
--- a/deployments/ci-cd.mdx
+++ b/deployments/ci-cd.mdx
@@ -10,7 +10,7 @@ This guide sets up a CI/CD pipeline using GitHub Actions and Cerebrium's Service
{" "}
Maintaining **separate development and production apps** in separate projects
- is recommended, so that changes can be safely tested before going live.
+ is recommended, so that you can safely test changes before going live.
### 1. Authenticating to Cerebrium
@@ -23,7 +23,7 @@ The GitHub Action workflow uses a **Service Account key** to authenticate to Cer

-### 2. Define secrets in a GitHub environment
+### 2. Define Secrets in a GitHub Environment
Store this key in a secret for use in GitHub Actions workflows.
diff --git a/deployments/multi-region-deployment.mdx b/deployments/multi-region-deployment.mdx
index 657c823b..de03f8bd 100644
--- a/deployments/multi-region-deployment.mdx
+++ b/deployments/multi-region-deployment.mdx
@@ -6,9 +6,9 @@ description: Run a Cerebrium app globally across multiple regions for more GPU c
Deploy an app once and run it in multiple regions. The `region` parameter in `cerebrium.toml` controls placement: run globally on whatever capacity is available (recommended), or pin the app to a specific region. The parameter is optional; when omitted, the platform chooses placement automatically based on the app's hardware requirements.
- Multi-region deployment is currently in **beta**. Rapid updates and
- improvements will be made over the next few months to bring full functionality
- to life. Please reach out on our [Discord](https://discord.gg/ATj6USmeE2)
+ Multi-region deployment is currently in **beta**. We will make rapid updates
+ and improvements over the next few months to bring full functionality to
+ life. Please reach out on our [Discord](https://discord.gg/ATj6USmeE2)
about features/functionality you would like to see.
@@ -125,7 +125,7 @@ The output includes the region of each container.
## Storage
-Persistent storage is managed per region: an app has an independent `/persistent-storage` volume in each region it runs in, however placement is configured. Files written in one region are not guaranteed to be available in other regions, and region-local caches, such as model weights downloaded on first load, fill independently per region.
+Persistent storage is managed per region: an app has an independent `/persistent-storage` volume in each region it runs in, however placement is configured. Files written in one region are not guaranteed to be available in other regions. Region-local caches, such as model weights downloaded on first load, fill independently per region.
Apps deployed with `region = "global"` also mount `/global-persistent-storage`, a single volume shared across every region the app runs in. Files written there are visible from all regions, and reads are cached per region. Use the global volume for data that must be available everywhere and `/persistent-storage` for region-local data. Manage files on the global volume by passing `--region global` to the file commands. See [Managing Files](/storage/managing-files#global-storage).
diff --git a/endpoints/async.mdx b/endpoints/async.mdx
index 98c2b22c..46a0bbcb 100644
--- a/endpoints/async.mdx
+++ b/endpoints/async.mdx
@@ -3,7 +3,7 @@ title: "Async requests"
description: Run Cerebrium functions asynchronously with the async query parameter, get a run_id back instantly, and forward results via a webhook endpoint.
---
-Some apps require asynchronous "fire-and-forget" execution. In this model, Cerebrium handles running the function, while the developer is responsible for ensuring data leaves the function (e.g. via a webhook).
+Some apps require asynchronous "fire-and-forget" execution. In this model, Cerebrium handles running the function, while you are responsible for ensuring data leaves the function (e.g. via a webhook).
Enable async execution by adding the `async=true` query parameter to the request:
@@ -30,9 +30,9 @@ X-Request-Id: 21eb3b98-4b10-9ad6-8681-a47172828024
{"run_id":"21eb3b98-4b10-9ad6-8681-a47172828024"}
```
-Async functions run for a maximum of **12 hours**, bounded by the `response_grace_period` in `cerebrium.toml`. This defaults to 15 minutes — update it to match the maximum time the task needs.
+Async functions run for a maximum of **12 hours**, bounded by the `response_grace_period` in `cerebrium.toml`. This defaults to 15 minutes. Update it to match the maximum time the task needs.
-Cerebrium runs the HTTP request in the background, but the function itself must still behave **synchronously** — it must complete its work and return a result.
+Cerebrium runs the HTTP request in the background, but the function itself must still behave **synchronously**. It must complete its work and return a result.
Returning a response while the application is still processing causes Cerebrium to begin terminating the container. Only return once all processing is finished.
Because async calls do not return a response to the caller, the function must export any relevant data itself. Combine async execution with a `webhookEndpoint` to have Cerebrium automatically forward the function's response body:
@@ -44,4 +44,4 @@ curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx//run?async=true&w
--data '{"param": "hello world"}'
```
-This is a proxy-level feature — no code changes are required to use webhook forwarding. In the dashboard, the function is marked **async** but still shows the status of the internal synchronous call (e.g. if the call failed, the async request state is `failure`).
+This is a proxy-level feature. No code changes are required to use webhook forwarding. In the dashboard, the function is marked **async** but still shows the status of the internal synchronous call (e.g. if the call failed, the async request state is `failure`).
diff --git a/endpoints/inference-api.mdx b/endpoints/inference-api.mdx
index bc854968..4f338af8 100644
--- a/endpoints/inference-api.mdx
+++ b/endpoints/inference-api.mdx
@@ -3,11 +3,11 @@ title: "REST API"
description: Call your Cerebrium apps over the REST API with POST requests and JWT authentication, and understand response formats and HTTP status codes.
---
-All functions on Cerebrium are accessible via POST requests, unless marked private by prefixing the function name with an underscore (e.g. `_private_function()`). Authenticate using the JWT token from the **API Keys** section of the dashboard. Endpoints require this token only when `cerebrium.toml` sets [`disable_auth = false`](/toml-reference/toml-reference) — authentication is disabled by default.
+All functions on Cerebrium are accessible via POST requests, unless marked private by prefixing the function name with an underscore (e.g. `_private_function()`). Authenticate using the JWT token from the **API Keys** section of the dashboard. Endpoints require this token only when `cerebrium.toml` sets [`disable_auth = false`](/toml-reference/toml-reference). Authentication is disabled by default.
## Request format
-The POST request follows the structure below, where `{function}` is the name of the function to invoke. In this example, `predict()` from `main.py` is called.
+The POST request follows the structure below, where `{function}` is the name of the function to invoke. This example calls `predict()` from `main.py`.
```bash
curl --location --request POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/{function}' \
diff --git a/endpoints/streaming.mdx b/endpoints/streaming.mdx
index 9f6bd85f..c770355c 100644
--- a/endpoints/streaming.mdx
+++ b/endpoints/streaming.mdx
@@ -7,7 +7,7 @@ Streaming sends live output from a model over a server-sent event (SSE) stream.
It works with any Python object that implements the iterator or generator protocol.
The generator/iterator must `yield` data, which is sent downstream via the `text/event-stream` Content-Type.
-Data can be sent in JSON format and decoded on the client side.
+You can send data in JSON format and decode it on the client side.
A minimal example:
diff --git a/endpoints/webhook.mdx b/endpoints/webhook.mdx
index 6f38c879..3125379b 100644
--- a/endpoints/webhook.mdx
+++ b/endpoints/webhook.mdx
@@ -12,11 +12,11 @@ curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx//run?webhookEndpo
--data '{"param": "hello world"}'
```
-The proxy forwards the response body as a POST request to the specified webhook — no code changes required. Ensure the webhook endpoint accepts POST requests. Webhook forwarding works with both Cortex and Custom runtimes, but only for HTTP requests (not WebSockets).
+The proxy forwards the response body as a POST request to the specified webhook. No code changes are required. Ensure the webhook endpoint accepts POST requests. Webhook forwarding works with both Cortex and Custom runtimes, but only for HTTP requests (not WebSockets).
## Retry Behavior
-Webhook requests are sent asynchronously and do not block the function's response. Failed deliveries are retried automatically:
+Webhook requests are sent asynchronously and do not block the function's response. Cerebrium retries failed deliveries automatically:
- **Maximum attempts**: 3 attempts total
- **Delay between retries**: Up to 5 seconds between attempts
@@ -33,7 +33,7 @@ If all attempts fail, the error is logged but will not affect your function's re
**The webhook endpoint should be idempotent.** A webhook may be retried even
- if it was already processed — for example, if the Cerebrium backend does not
+ if it was already processed. For example, the Cerebrium backend may not
receive the success response due to network issues. Design the endpoint to
handle duplicate deliveries gracefully.
diff --git a/endpoints/websockets.mdx b/endpoints/websockets.mdx
index 8950f47d..1eb5cd64 100644
--- a/endpoints/websockets.mdx
+++ b/endpoints/websockets.mdx
@@ -38,7 +38,7 @@ Test the WebSocket endpoint using websocat, a command-line WebSocket client:
websocat wss://api.cerebrium.ai/v4/p-xxxxxxxx//
```
-## Implementing the WebSocket Endpoint
+## Implementing the WebSocket endpoint
Example WebSocket endpoint using FastAPI:
@@ -55,7 +55,7 @@ async def websocket_endpoint(websocket: WebSocket):
await websocket.close()
```
-## Additional Info
+## Additional info
Client-side Implementation: Handle the WebSocket connection properly on the client, including error handling and reconnection logic.
diff --git a/hardware/cpu-and-memory.mdx b/hardware/cpu-and-memory.mdx
index 1a52b9df..5ebfeb6a 100644
--- a/hardware/cpu-and-memory.mdx
+++ b/hardware/cpu-and-memory.mdx
@@ -32,8 +32,8 @@ memory = 16.0 # Memory in GB
Allocate system memory equal to the GPU's VRAM capacity as a baseline. This accounts for initial model loading and compilation before GPU transfer. Applications terminate with an Out of Memory (OOM) error if they exceed the specified memory limit.
- Memory and CPU are billed based on usage, which reduces costs for end-users
- and doesn’t require the overprovisioning of an entire instance.
+ Memory and CPU are billed based on usage, which reduces your costs and
+ doesn’t require the overprovisioning of an entire instance.
## Resource Limits
@@ -60,4 +60,4 @@ The Transformers library provides memory optimization through the `low_cpu_mem_u
## Resource Monitoring
-The platform monitors CPU utilization and throttling events to identify performance bottlenecks. Memory usage and OOM events are tracked to prevent application failures.
+The platform monitors CPU utilization and throttling events to identify performance bottlenecks. The platform tracks memory usage and OOM events to prevent application failures.
diff --git a/hardware/using-cuda.mdx b/hardware/using-cuda.mdx
index 9363c521..a6b9b37d 100644
--- a/hardware/using-cuda.mdx
+++ b/hardware/using-cuda.mdx
@@ -38,7 +38,7 @@ docker_base_image_url = "nvidia/cuda:12.1.1-runtime-ubuntu22.04"
## Cold-Start Optimization
-Image size and complexity directly impact cold-start performance — the time needed to initialize an app from an inactive state. Cerebrium uses a content-addressable file system that selectively pulls only required files, but larger images still affect startup times.
+Image size and complexity directly impact cold-start performance: the time needed to initialize an app from an inactive state. Cerebrium uses a content-addressable file system that selectively pulls only required files, but larger images still affect startup times.
### Image Size Considerations
diff --git a/hardware/using-gpus.mdx b/hardware/using-gpus.mdx
index b6ad0c9c..ed60069f 100644
--- a/hardware/using-gpus.mdx
+++ b/hardware/using-gpus.mdx
@@ -36,7 +36,7 @@ The platform offers GPUs ranging from cost-effective development options to high
model generation and model name to avoid ambiguity.
-### Plan availability
+### Plan Availability
Compute types are gated by plan:
@@ -53,7 +53,7 @@ Deploying with a compute type outside the project's plan is rejected at deploy t
## Multi-GPU Configuration
-Multiple GPUs are configured in the `cerebrium.toml` file:
+Configure multiple GPUs in the `cerebrium.toml` file:
```toml
[cerebrium.hardware]
diff --git a/networking/custom-domains.mdx b/networking/custom-domains.mdx
index 714aa36d..42f1a43a 100644
--- a/networking/custom-domains.mdx
+++ b/networking/custom-domains.mdx
@@ -10,7 +10,7 @@ Once configured, API calls use the custom domain while keeping the same path str
- Support for apex domains (`example.com`) and subdomains (`api.example.com`)
- Automatic SSL certificate provisioning and renewal
-- Project-level domains - one domain serves all apps in a project
+- Project-level domains: one domain serves all apps in a project
- Multiple domains can point to the same project
- Professional branding with custom domains instead of `*.cerebrium.ai` URLs
@@ -32,7 +32,7 @@ Once configured, API calls use the custom domain while keeping the same path str
### Step 2: Configure DNS Records
-After creating the domain, DNS configuration instructions will be displayed. Create a CNAME record at the DNS provider using the DNS record Cerebrium generates.
+After you create the domain, the dashboard displays DNS configuration instructions. Create a CNAME record at the DNS provider using the DNS record Cerebrium generates.
DNS record details can also be found later by clicking "DNS Record" in the
@@ -78,7 +78,7 @@ After creating the domain, DNS configuration instructions will be displayed. Cre
2. Cerebrium will automatically attempt to validate DNS records every 30 minutes for up to 2 days
3. To trigger an immediate validation attempt, click "Validation Status" -> "Validate Domain" (this works even after the 2-day window has elapsed)
4. If validation fails, the dialog will show the last known error
-5. Once validated, SSL certificates will be automatically provisioned
+5. Once validated, Cerebrium automatically provisions SSL certificates
### Step 4: Start Using the Custom Domain
@@ -110,7 +110,7 @@ Cerebrium attempts to validate domains once every 30 minutes for up to 2 days. I
- **Pending**: Domain is waiting for DNS validation
- **Validated**: Domain is successfully validated and ready to use
-- **Failed**: DNS validation failed - click "Validation Status" to see the specific error
+- **Failed**: DNS validation failed. Click "Validation Status" to see the specific error
### Validation Errors
@@ -124,7 +124,7 @@ If a domain shows "Failed" status, the DNS record has been misconfigured. Common
- **Cloudflare**: Disable proxy (set to "DNS only" - gray cloud icon)
- **Route53**: Use simple routing policy, not weighted or latency-based
- **Namecheap**: Use "@" for apex domains, not "www" or blank
-- **GoDaddy**: CNAME records cannot be used with apex domains - consider using a subdomain or an ALIAS record
+- **GoDaddy**: CNAME records cannot be used with apex domains. Consider using a subdomain or an ALIAS record
### Common DNS Mistakes
diff --git a/other-topics/request-response-logging.mdx b/other-topics/request-response-logging.mdx
index 93f585cb..fefd1049 100644
--- a/other-topics/request-response-logging.mdx
+++ b/other-topics/request-response-logging.mdx
@@ -9,8 +9,8 @@ By default, the Cortex runtime logs all requests and responses. These logs appea
Two settings control logging behavior in the Cortex runtime:
-- `DISABLE_REQUEST_LOGS` - When enabled, prevents logging of incoming request data
-- `DISABLE_RESPONSE_LOGS` - When enabled, prevents logging of response data from your app
+- `DISABLE_REQUEST_LOGS`: When enabled, prevents logging of incoming request data
+- `DISABLE_RESPONSE_LOGS`: When enabled, prevents logging of response data from your app
These settings only affect the default Cortex runtime. If you are using a
diff --git a/other-topics/using-secrets.mdx b/other-topics/using-secrets.mdx
index b8356013..671a6e0e 100644
--- a/other-topics/using-secrets.mdx
+++ b/other-topics/using-secrets.mdx
@@ -5,7 +5,7 @@ description: Store API keys and credentials as encrypted secrets in Cerebrium, e
Secrets store API keys, passwords, and other sensitive information outside of code. Secrets are encrypted at rest (256-bit AES) and decrypted only at runtime.
-Secrets can be managed at both project and app levels. Project-level secrets are shared across all apps in your project, while app-level secrets are specific to an individual app. App secrets take precedence over project-wide secrets.
+You can manage secrets at both project and app levels. Project-level secrets are shared across all apps in your project, while app-level secrets are specific to an individual app. App secrets take precedence over project-wide secrets.
Each secret is exposed to the app as an environment variable.
@@ -31,7 +31,7 @@ def predict(run_id):
### Managing Secrets
-Secrets are created, updated, and deleted in your dashboard.
+You create, update, and delete secrets in your dashboard.

diff --git a/partner-services/deepgram.mdx b/partner-services/deepgram.mdx
index 985fafa5..9d3e7cd5 100644
--- a/partner-services/deepgram.mdx
+++ b/partner-services/deepgram.mdx
@@ -20,7 +20,7 @@ Consult the Deepgram representative on how to achieve parity with the Deepgram A
Deployments currently run Deepgram self-hosted release `260728`. Model files
- and `engine.toml`/`api.toml` settings should match that release — check with
+ and `engine.toml`/`api.toml` settings should match that release. Check with
your Deepgram Account Representative if you are unsure whether your model
files are current.
@@ -328,12 +328,12 @@ wget https://dpgr.am/bueller.wav
curl -X POST --data-binary @bueller.wav "https://api.cerebrium.ai/v4/p-xxxxxxxx/deepgram/v1/listen?model=nova-3&smart_format=true"
```
-Parameters accepted by the Deepgram service can be found in the [speech-to-text API reference](https://developers.deepgram.com/reference/speech-to-text-api/listen-streaming).
+You can find the parameters accepted by the Deepgram service in the [speech-to-text API reference](https://developers.deepgram.com/reference/speech-to-text-api/listen-streaming).
If 'disable_auth' in cerebrium.toml is set to false, include the inference
token in the Authorization header to authenticate with the Cerebrium service.
- The Deepgram API key is pulled automatically from secrets.
+ Cerebrium pulls the Deepgram API key automatically from secrets.
## API Key Configuration
@@ -361,6 +361,6 @@ Adjust these parameters based on traffic patterns and latency requirements.
## Usage Examples
-Cerebrium runs both Deepgram STT models and applications on the same network alongside LiveKit workers, reducing latency by approximately 400ms—a significant advantage for voice agent applications.
+Cerebrium runs both Deepgram STT models and applications on the same network alongside LiveKit workers, reducing latency by approximately 400ms. This is a significant advantage for voice agent applications.
For a complete implementation reference, see the [LiveKit Outbound Agent example](/v4/examples/livekit-outbound-agent).
diff --git a/partner-services/index.mdx b/partner-services/index.mdx
index e8415b61..63b9323e 100644
--- a/partner-services/index.mdx
+++ b/partner-services/index.mdx
@@ -16,8 +16,8 @@ Cerebrium offers specialized services in partnership with leading AI companies,
Available Partner Services:
-- [Deepgram](/partner-services/deepgram) - Speech-to-text (STT) services
-- [Rime](/partner-services/rime) - Text-to-speech (TTS) services
+- [Deepgram](/partner-services/deepgram): Speech-to-text (STT) services
+- [Rime](/partner-services/rime): Text-to-speech (TTS) services
## Benefits of Partner Services
diff --git a/partner-services/rime.mdx b/partner-services/rime.mdx
index f6e426d2..fa33cf33 100644
--- a/partner-services/rime.mdx
+++ b/partner-services/rime.mdx
@@ -56,7 +56,7 @@ replica_concurrency = 50
The Rime Server validates the API key directly.
-4. Run `cerebrium deploy` to deploy the Rime service - the output of which should appear as follows:
+4. Run `cerebrium deploy` to deploy the Rime service. The output should appear as follows:
```
App Dashboard: https://dashboard.cerebrium.ai/projects/p-xxxxxxxx/apps/p-xxxxxxxx-rime
diff --git a/scaling/batching-concurrency.mdx b/scaling/batching-concurrency.mdx
index 776a82a0..fe0c6c34 100644
--- a/scaling/batching-concurrency.mdx
+++ b/scaling/batching-concurrency.mdx
@@ -74,6 +74,6 @@ fastapi = "latest"
for more information.
-Custom batching provides full control over request grouping and processing — particularly useful for frameworks without native batching support. The [Container Images Guide](/container-images/defining-container-images#custom-runtimes) provides detailed implementation instructions.
+Custom batching provides full control over request grouping and processing. This is particularly useful for frameworks without native batching support. The [Container Images Guide](/container-images/defining-container-images#custom-runtimes) provides detailed implementation instructions.
Concurrency enables parallel request handling; batching optimizes how those requests are processed. Together, they improve resource utilization and throughput.
diff --git a/scaling/graceful-termination.mdx b/scaling/graceful-termination.mdx
index c569ce16..12892e5c 100644
--- a/scaling/graceful-termination.mdx
+++ b/scaling/graceful-termination.mdx
@@ -5,7 +5,7 @@ description: Handle SIGTERM signals in custom runtimes on Cerebrium to finish in
## Graceful Termination
-Cerebrium runs in a shared, multi-tenant environment. The platform continuously adjusts capacity — spinning down nodes and launching new ones to scale, optimize compute usage, and roll out updates. Workloads are migrated to new nodes during this process. Applications also have metric-based autoscaling criteria that dictate when instances scale, remain active, or shift during deployments. Implement graceful termination to prevent requests from ending prematurely when instances are marked for termination.
+Cerebrium runs in a shared, multi-tenant environment. The platform continuously adjusts capacity: spinning down nodes and launching new ones to scale, optimize compute usage, and roll out updates. The platform migrates workloads to new nodes during this process. Applications also have metric-based autoscaling criteria that dictate when instances scale, remain active, or shift during deployments. Implement graceful termination to prevent requests from ending prematurely when instances are marked for termination.
## Understanding Instance Termination
diff --git a/toml-reference/toml-reference.mdx b/toml-reference/toml-reference.mdx
index 16e73ed0..bff263fb 100644
--- a/toml-reference/toml-reference.mdx
+++ b/toml-reference/toml-reference.mdx
@@ -63,8 +63,8 @@ use_uv = true
Check your build logs for these indicators:
-- **UV_PIP_INSTALL_STARTED** - UV is successfully being used
-- **PIP_INSTALL_STARTED** - Standard pip installation (when `use_uv` is `false`)
+- **UV_PIP_INSTALL_STARTED**: UV is successfully being used
+- **PIP_INSTALL_STARTED**: Standard pip installation (when `use_uv` is `false`)
While UV is compatible with most packages, some edge cases may cause build
diff --git a/v4/examples/comfyUI.mdx b/v4/examples/comfyUI.mdx
index 0c97e303..310b1f3a 100644
--- a/v4/examples/comfyUI.mdx
+++ b/v4/examples/comfyUI.mdx
@@ -12,7 +12,7 @@ ComfyUI is a popular no-code interface for building complex stable diffusion wor
Production-scale deployment guidance for ComfyUI is limited. This tutorial covers deploying ComfyUI pipelines on Cerebrium as autoscaling API endpoints with pay-as-you-go compute. Find the full example code [here](https://github.com/CerebriumAI/examples/tree/master/7-image-and-video/1-comfyui).
-### Creating your Comfy UI workflow locally
+### Creating Your Comfy UI Workflow Locally
Create the workflow locally or using a rented GPU from [Lambda Labs](https://lambdalabs.com/). Ensure [ComfyUI is installed](https://github.com/comfyanonymous/ComfyUI#installing) in the local environment.
@@ -35,7 +35,7 @@ The example GitHub repository contains a workflow.json file. Click the “Load

-Don’t worry about the pre-filled values and prompts — they get overridden at inference time.
+Don’t worry about the pre-filled values and prompts. They get overridden at inference time.
To export the workflow in API format, click the gear icon (settings) in the top-right hovering panel. Ensure that “Enable dev mode” is selected, then close the popup.
@@ -61,7 +61,7 @@ Alter the workflow_api.json file to include placeholders for user values at infe
- Replace line 58, the input text of node 7 with: "\{\{negative_prompt\}\}"
- Replace line 108, the image of node 11 with: "\{\{controlnet_image\}\}"
-### FastAPI app
+### FastAPI App
In addition to the `main.py` file below (which runs the FastAPI server and initializes ComfyUI on application start), a separate `helpers.py` file contains utility functions for working with ComfyUI. Find the helper code [here](https://github.com/CerebriumAI/examples/blob/master/7-image-and-video/1-comfyui/helpers.py). Create a file named `helpers.py` and copy the code into it.
diff --git a/v4/examples/deploy-a-vision-language-model-with-sglang.mdx b/v4/examples/deploy-a-vision-language-model-with-sglang.mdx
index 6b6a721f..732ea24b 100644
--- a/v4/examples/deploy-a-vision-language-model-with-sglang.mdx
+++ b/v4/examples/deploy-a-vision-language-model-with-sglang.mdx
@@ -7,7 +7,7 @@ This tutorial deploys a Vision Language Model (VLM) using SGLang on Cerebrium. A
The example builds an intelligent ad analysis system that evaluates advertisements across multiple dimensions, scoring how the advertisement relates to the business in question and how it performs on the given criteria.
-SGLang (Structured Generation Language) differs from other inference frameworks such as vLLM and TensorRT by focusing on structured generation and complex multi-step LLM workflows. SGLang is being used in production by teams at xAI and Deepseek to power their core language model capabilities making it a trusted choice.
+SGLang (Structured Generation Language) differs from other inference frameworks such as vLLM and TensorRT by focusing on structured generation and complex multi-step LLM workflows. Teams at xAI and Deepseek use SGLang in production to power their core language model capabilities, making it a trusted choice.
### SGLang Architecture
@@ -109,7 +109,7 @@ entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
### Step 3: Implement the Ad Analysis Logic
-Cerebrium does not enforce any special class design or application architecture — write Python code as if running locally. The code below sets up the SGLang Runtime Engine (Backend) with FastAPI and loads the model on container startup. The first request incurs a model load, but subsequent requests execute instantaneously.
+Cerebrium does not enforce any special class design or application architecture. Write Python code as if running locally. The code below sets up the SGLang Runtime Engine (Backend) with FastAPI and loads the model on container startup. The first request incurs a model load, but subsequent requests execute instantaneously.
In your `main.py` file:
diff --git a/v4/examples/deploy-an-llm-with-tensorrtllm-tritonserver.mdx b/v4/examples/deploy-an-llm-with-tensorrtllm-tritonserver.mdx
index e8c57f01..f7d41977 100644
--- a/v4/examples/deploy-an-llm-with-tensorrtllm-tritonserver.mdx
+++ b/v4/examples/deploy-an-llm-with-tensorrtllm-tritonserver.mdx
@@ -15,7 +15,7 @@ You can view the final implementation [here](https://github.com/CerebriumAI/exam
NVIDIA TensorRT is a software development kit for high-performance deep learning inference. It compiles model weights into optimized engines that run more efficiently on specific GPU hardware through CUDA-level optimizations, custom kernels, and optional quantization.
-TensorRT requires you to specify optimization parameters upfront - GPU architecture, batch size, precision (FP8, INT8, etc.), and input/output shapes. This specialization allows TensorRT to generate highly optimized inference engines that maximize GPU utilization, reduce latency, and lower inference costs compared to serving raw model weights.
+TensorRT requires you to specify optimization parameters upfront: GPU architecture, batch size, precision (FP8, INT8, etc.), and input/output shapes. This specialization allows TensorRT to generate highly optimized inference engines that maximize GPU utilization, reduce latency, and lower inference costs compared to serving raw model weights.
### Why Triton?
@@ -57,7 +57,7 @@ To download the model, [request access](https://huggingface.co/meta-llama/Llama-
## Implementation
-All files should be placed in the same project directory.
+Place all files in the same project directory.
### Triton Model Configuration
diff --git a/v4/examples/gpt-oss.mdx b/v4/examples/gpt-oss.mdx
index 4b39eed2..aeb423b0 100644
--- a/v4/examples/gpt-oss.mdx
+++ b/v4/examples/gpt-oss.mdx
@@ -3,7 +3,7 @@ title: "Serving GPT-OSS with vLLM"
description: Serve OpenAI GPT-OSS open weight models with vLLM on Cerebrium, covering MoE architecture, MXFP4 quantization and H100 GPU deployment setup.
---
-GPT recently released GPT-OSS ([gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) and [gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)), two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.0 license, these models outperform similarly sized open models on reasoning tasks, demonstrate strong tool use capabilities, and are optimized for efficient deployment on consumer hardware.
+GPT recently released GPT-OSS ([gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) and [gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)), two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.0 license, these models outperform similarly sized open models on reasoning tasks and demonstrate strong tool use capabilities. They are also optimized for efficient deployment on consumer hardware.
## What Makes GPT-OSS Special?
@@ -70,7 +70,7 @@ Key configuration details:
### Deploy & Test
-Deploy by running `cerebrium deploy`. The environment is created and the model is downloaded.
+Deploy by running `cerebrium deploy`. Cerebrium creates the environment and downloads the model.
Test the endpoint with the following request:
diff --git a/v4/examples/langchain-langsmith.mdx b/v4/examples/langchain-langsmith.mdx
index 3cbc0136..7b249606 100644
--- a/v4/examples/langchain-langsmith.mdx
+++ b/v4/examples/langchain-langsmith.mdx
@@ -9,7 +9,7 @@ You can find the final version of the code [here](https://github.com/CerebriumAI
### Concepts
-This app requires calendar interaction based on user instructions — an ideal use case for an agent with function (tool) calling capabilities. LangChain provides extensive agent support, and its companion tool LangSmith makes monitoring integration straightforward.
+This app requires calendar interaction based on user instructions: an ideal use case for an agent with function (tool) calling capabilities. LangChain provides extensive agent support, and its companion tool LangSmith makes monitoring integration straightforward.
A tool refers to any framework, utility, or system with defined functionality for specific use cases, such as searching Google or retrieving credit card transactions.
@@ -47,7 +47,7 @@ agent_executor.invoke({"input": "what's 3 plus 5 raised to the 2.743. also what'
### Setup Cal.com
-[Cal.com](https://cal.com) provides the calendar management foundation. Create an account [here](https://app.cal.com/signup) if needed. Cal serves as the source of truth — updates to time zones or working hours automatically reflect in the assistant's responses.
+[Cal.com](https://cal.com) provides the calendar management foundation. Create an account [here](https://app.cal.com/signup) if needed. Cal serves as the source of truth. Updates to time zones or working hours automatically reflect in the assistant's responses.
After creating your account:
@@ -127,7 +127,7 @@ The API key is now confirmed working and pulling calendar information. The API c
- **/availability**: Get your availability
- **/bookings**: Book a slot
-### Cerebrium setup
+### Cerebrium Setup
Set up Cerebrium:
@@ -281,7 +281,7 @@ The agent executor consists of:
- Defines the agent’s role, goals, and situational behavior. More precise instructions yield better results.
- Chat History stores previous messages for conversation context.
- Input receives new input from the end user.
-- The GPT-3.5 model serves as the LLM. Swap to Anthropic or any other provider by replacing this one line — LangChain makes this seamless.
+- The GPT-3.5 model serves as the LLM. Swap to Anthropic or any other provider by replacing this one line. LangChain makes this seamless.
- Finally, these components combine with the tools to create an agent executor.
### Setup Chatbot
@@ -334,7 +334,7 @@ if __name__ == "__main__":
This code:
-- Defines a Pydantic object specifying the expected API parameters — user prompt and session ID.
+- Defines a Pydantic object specifying the expected API parameters: user prompt and session ID.
- The predict function (Cerebrium’s API entry point) passes the prompt and session ID to the agent and returns results.
diff --git a/v4/examples/livekit-outbound-agent.mdx b/v4/examples/livekit-outbound-agent.mdx
index 699c8d14..f22e4450 100644
--- a/v4/examples/livekit-outbound-agent.mdx
+++ b/v4/examples/livekit-outbound-agent.mdx
@@ -7,11 +7,11 @@ Voice agents are transforming business operations by introducing efficiencies an
Outbound agents handle tasks like appointment scheduling, lead qualification, and customer follow-ups while maintaining a human-like, conversational tone. They replace manual efforts with intelligent automation, improving customer engagement in sales, healthcare, hospitality, and collections.
-This tutorial sets up an outbound calling agent that does a warm transfer — handing off the call to a real person using LiveKit. A fitting example: a support center where an agent collects data before connecting to a real person.
+This tutorial sets up an outbound calling agent that does a warm transfer: handing off the call to a real person using LiveKit. A fitting example: a support center where an agent collects data before connecting to a real person.
-The final code implementation and complete example can be found in our examples Github repository [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/7-outbound-livekit-agent).
+The final code implementation and complete example can be found in our examples GitHub repository [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/7-outbound-livekit-agent).
-### Cerebrium setup
+### Cerebrium Setup
Install the Cerebrium CLI and create a project:
@@ -82,7 +82,7 @@ Setting up an outbound calling agent requires a SIP trunk in Twilio. A SIP trunk
To secure the SIP trunk, create a credential list in the Twilio console dashboard. Navigate to “Voice”, then “Manage”, then “Credential lists”.
- Click the plus icon and add a friendly name, username, and password. Save these credentials — they are required in a later step.
+ Click the plus icon and add a friendly name, username, and password. Save these credentials. They are required in a later step.

@@ -113,7 +113,7 @@ Run the following in your CLI: `lk sip outbound create outbound-trunk.json`
The output returns the trunk ID, which is required in a later step.
-### Services setup
+### Services Setup
The following services power the outbound AI agent:
@@ -423,9 +423,9 @@ Run the following command:
cerebrium deploy
```
-This installs all necessary packages and deploys the application. LiveKit workers run in the cloud — test by running the `test.py` script locally and checking logs on the Cerebrium dashboard.
+This installs all necessary packages and deploys the application. LiveKit workers run in the cloud. Test by running the `test.py` script locally and checking logs on the Cerebrium dashboard.
-### Further Improvements:
+### Further Improvements
To reduce latency, Cerebrium partners with Deepgram and Rime to run STT and TTS models locally alongside the LiveKit worker, reducing latency by ~400ms.
diff --git a/v4/examples/openai-compatible-endpoint-vllm.mdx b/v4/examples/openai-compatible-endpoint-vllm.mdx
index 04ea4fd6..fe833231 100644
--- a/v4/examples/openai-compatible-endpoint-vllm.mdx
+++ b/v4/examples/openai-compatible-endpoint-vllm.mdx
@@ -7,7 +7,7 @@ This tutorial creates an OpenAI-compatible endpoint that works with any open-sou
To see the final code implementation, you can view it [here](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/1-openai-compatible-endpoint)
-### Cerebrium setup
+### Cerebrium Setup
Create a Cerebrium account by signing up [here](https://dashboard.cerebrium.ai/register) and follow the [installation docs](https://docs.cerebrium.ai/getting-started/installation).
diff --git a/v4/examples/realtime-voice-agents.mdx b/v4/examples/realtime-voice-agents.mdx
index 602acf48..d86fcf11 100644
--- a/v4/examples/realtime-voice-agents.mdx
+++ b/v4/examples/realtime-voice-agents.mdx
@@ -14,7 +14,7 @@ The application has 3–4 parts:
- A Deepgram TTS/STT service (requires a Deepgram Enterprise account)
- A self-hosted LLM using the vLLM framework
-Low latency is achieved because each service is hosted within Cerebrium — communication across containers is less than 10ms with no network latency overhead.
+Low latency is achieved because each service is hosted within Cerebrium. Communication across containers is less than 10ms with no network latency overhead.

@@ -22,7 +22,7 @@ You can find the final version of the code [here](https://github.com/CerebriumAI
Create a Cerebrium account by signing up [here](https://dashboard.cerebrium.ai/register) and follow the [installation docs](https://docs.cerebrium.ai/getting-started/installation).
-### Deepgram deployment
+### Deepgram Deployment
See the [Partner Services page](/partner-services/deepgram) to deploy a Deepgram service on Cerebrium.
@@ -60,7 +60,7 @@ vllm = "latest"
pydantic = "latest"
```
-Add the following code to `main.py` — this uses the vLLM framework and makes it OpenAI compatible:
+Add the following code to `main.py`. This uses the vLLM framework and makes it OpenAI compatible:
```
import os
@@ -162,7 +162,7 @@ Run `cerebrium deploy` to make it live. The deployment URL appears in the dashbo
Adjust the GPU hardware and `replica_concurrency` in `cerebrium.toml` to control how many concurrent calls the LLM handles.
-### Pipecat setup
+### Pipecat Setup
Run the following command to create the pipecat-agent: `cerebrium init pipecat-agent`. The [Pipecat framework](https://docs.pipecat.ai/getting-started/overview) orchestrates the services to create a voice agent.
@@ -460,7 +460,7 @@ Deploy to Cerebrium by running `cerebrium deploy`.
The endpoints are used in the frontend interface below.
-## Connect frontend
+## Connect Frontend
A public fork of the PipeCat frontend demonstrates this application. Clone the repo [here](https://github.com/CerebriumAI/web-client-ui).
diff --git a/v4/examples/sdxl.mdx b/v4/examples/sdxl.mdx
index 530a1e86..5563fefc 100644
--- a/v4/examples/sdxl.mdx
+++ b/v4/examples/sdxl.mdx
@@ -83,7 +83,7 @@ class Item(BaseModel):
The code uses Pydantic for data validation. The `prompt` and `url` parameters are required; all others are optional. Missing required parameters trigger an automatic error message.
-## Instantiate model
+## Instantiate Model
The SDXL model loads outside the `predict` function since it only needs to load once at startup. The model downloads during initial deployment and is automatically cached in persistent storage for subsequent use.
diff --git a/v4/examples/transcribe-whisper.mdx b/v4/examples/transcribe-whisper.mdx
index 4ca9999d..76be1f44 100644
--- a/v4/examples/transcribe-whisper.mdx
+++ b/v4/examples/transcribe-whisper.mdx
@@ -3,7 +3,7 @@ title: "Transcribe 1 hour podcast"
description: Transcribe hour long podcasts and audio files with Distil Whisper on Cerebrium using base64 uploads or file URLs and webhooks for long jobs
---
-This tutorial transcribes an hour-long audio file using Distill Whisper — an optimized version of Whisper-large-v2 that's 60% faster while maintaining accuracy within 1% of the original. The endpoint accepts either a base64-encoded string of the audio file or a URL to download the audio file.
+This tutorial transcribes an hour-long audio file using Distill Whisper: an optimized version of Whisper-large-v2 that's 60% faster while maintaining accuracy within 1% of the original. The endpoint accepts either a base64-encoded string of the audio file or a URL to download the audio file.
To see the final implementation, you can view it [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/1-whisper-transcription)
@@ -27,7 +27,7 @@ openai-whisper = "latest"
pydantic = "latest"
```
-Create a `util.py` file for utility functions — downloading a file from a URL or converting a base64 string to a file:
+Create a `util.py` file for utility functions that download a file from a URL or convert a base64 string to a file:
```python
import base64
@@ -66,11 +66,11 @@ class Item(BaseModel):
webhook_endpoint: Optional[HttpUrl]
```
-Pydantic handles data validation. While `audio` and `file_url` are optional parameters, at least one must be provided. The `webhook_endpoint` parameter, automatically included by Cerebrium in every request, is useful for long-running requests.
+Pydantic handles data validation. While `audio` and `file_url` are optional parameters, you must provide at least one. The `webhook_endpoint` parameter, automatically included by Cerebrium in every request, is useful for long-running requests.
-Note: Cerebrium has a 3-minute timeout for each inference request. For long audio files (2+ hours) that take several minutes to process, use a `webhook_endpoint` — a URL where Cerebrium sends a POST request with the function's results.
+Note: Cerebrium has a 3-minute timeout for each inference request. For long audio files (2+ hours) that take several minutes to process, use a `webhook_endpoint`: a URL where Cerebrium sends a POST request with the function's results.
-## Setup Model and inference
+## Setup Model and Inference
Import the required packages and load the Whisper model. The model downloads during initial deployment and is automatically cached in persistent storage for subsequent use. Loading the model outside the `predict` function ensures this code only runs on cold start (startup). For warm containers, only the `predict` function executes for inference.
@@ -152,7 +152,7 @@ curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/1-whisper-transcription/
--data '{"file_url": "https://your-public-url.com/test.mp3"}'
```
-The response returns immediately with a 202 status code and a `run_id` — a unique identifier to correlate the result with the initial workload.
+The response returns immediately with a 202 status code and a `run_id`: a unique identifier to correlate the result with the initial workload.
The endpoint returns results in this format:
diff --git a/v4/examples/twilio-voice-agent.mdx b/v4/examples/twilio-voice-agent.mdx
index 778b2f4b..e14c95d0 100644
--- a/v4/examples/twilio-voice-agent.mdx
+++ b/v4/examples/twilio-voice-agent.mdx
@@ -9,7 +9,7 @@ The example uses [PipeCat](https://www.pipecat.ai/) to handle component integrat
You can find the final version of the code [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/4-twilio-voice-agent)
-### Cerebrium setup
+### Cerebrium Setup
Set up Cerebrium:
@@ -107,7 +107,7 @@ healthcheck_endpoint = "/health"
You can read more about running custom web servers [here](/container-images/custom-web-servers).
-### Twilio setup
+### Twilio Setup
Twilio provides cloud communications APIs for messaging, voice, video, and authentication. Other providers work as well. Sign up for a free account [here](https://www.twilio.com/try-twilio).
@@ -246,7 +246,7 @@ A successful deployment looks like this:

-Test by calling the Twilio number — the agent responds automatically.
+Test by calling the Twilio number. The agent responds automatically.
### Scaling Pipecat
diff --git a/v4/examples/wandb-sweep.mdx b/v4/examples/wandb-sweep.mdx
index 56324755..94cd1267 100644
--- a/v4/examples/wandb-sweep.mdx
+++ b/v4/examples/wandb-sweep.mdx
@@ -57,7 +57,7 @@ pip install wandb
wandb login
```
-A link prints in the terminal — click it and copy the API key back into the terminal.
+A link prints in the terminal. Click it and copy the API key back into the terminal.
Add the W&B API key to Cerebrium secrets. In the [Cerebrium Dashboard](https://dashboard.cerebrium.ai/), navigate to the “secrets” tab in the left sidebar. Add the following:
@@ -289,7 +289,7 @@ def train_model(params: Dict):
You can read a deeper explanation of the training script [here](https://www.datacamp.com/tutorial/fine-tuning-llama-3-2) but here's a high-level explanation of the code in bullet points:
- This code sets up a fine-tuning pipeline for a Large Language Model (specifically Llama 3.2) using several modern training techniques:
-- Takes a dictionary of parameters for flexible training configurations — the hyperparameter sweep.
+- Takes a dictionary of parameters for flexible training configurations: the hyperparameter sweep.
- Loads a customer support dataset from Hugging Face and formats it into chat template format
- Implements QLoRA (Quantized Low-Rank Adaptation) for efficient fine-tuning.
- Uses Weights & Biases (Wandb) for experiment tracking, logging results to the Wandb dashboard.
@@ -307,7 +307,7 @@ This command:
2. Deploys the training script as an endpoint
3. Returns a POST URL (save this for later)
-Cerebrium requires no special decorators or syntax — wrap the training code in a function. The endpoint automatically scales based on request volume, making it ideal for hyperparameter sweeps.
+Cerebrium requires no special decorators or syntax. Wrap the training code in a function. The endpoint automatically scales based on request volume, making it ideal for hyperparameter sweeps.
### Hyperparameter Sweep
From f3f464c42ecfd95cd72656485fc46ffc7654250c Mon Sep 17 00:00:00 2001
From: "mintlify[bot]"
Date: Wed, 26 Aug 2026 20:32:26 +0000
Subject: [PATCH 3/3] Prettified Code!
---
deployments/multi-region-deployment.mdx | 6 +++---
hardware/cpu-and-memory.mdx | 4 ++--
2 files changed, 5 insertions(+), 5 deletions(-)
diff --git a/deployments/multi-region-deployment.mdx b/deployments/multi-region-deployment.mdx
index de03f8bd..d67a775a 100644
--- a/deployments/multi-region-deployment.mdx
+++ b/deployments/multi-region-deployment.mdx
@@ -7,9 +7,9 @@ Deploy an app once and run it in multiple regions. The `region` parameter in `ce
Multi-region deployment is currently in **beta**. We will make rapid updates
- and improvements over the next few months to bring full functionality to
- life. Please reach out on our [Discord](https://discord.gg/ATj6USmeE2)
- about features/functionality you would like to see.
+ and improvements over the next few months to bring full functionality to life.
+ Please reach out on our [Discord](https://discord.gg/ATj6USmeE2) about
+ features/functionality you would like to see.
## Why Use Multi-Region Deployment
diff --git a/hardware/cpu-and-memory.mdx b/hardware/cpu-and-memory.mdx
index 5ebfeb6a..122d6342 100644
--- a/hardware/cpu-and-memory.mdx
+++ b/hardware/cpu-and-memory.mdx
@@ -32,8 +32,8 @@ memory = 16.0 # Memory in GB
Allocate system memory equal to the GPU's VRAM capacity as a baseline. This accounts for initial model loading and compilation before GPU transfer. Applications terminate with an Out of Memory (OOM) error if they exceed the specified memory limit.
- Memory and CPU are billed based on usage, which reduces your costs and
- doesn’t require the overprovisioning of an entire instance.
+ Memory and CPU are billed based on usage, which reduces your costs and doesn’t
+ require the overprovisioning of an entire instance.
## Resource Limits