Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions container-images/custom-dockerfiles.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -48,7 +48,7 @@ CMD ["python", "-m", "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8192

Dockerfiles for Cerebrium have three requirements:

1. Expose a port with the `EXPOSE` command - this port is referenced in `cerebrium.toml`
1. Expose a port with the `EXPOSE` command. This port is referenced in `cerebrium.toml`
2. Include a `CMD` command to specify the container's startup process (typically the server)
3. Set the working directory with `WORKDIR` to ensure correct file paths (defaults to root if not specified)

Expand Down Expand Up @@ -87,7 +87,7 @@ entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8192"]

## Building Generic Dockerized Apps

Cerebrium supports non-Python apps as long as a Dockerfile is provided. The following example shows a Rust-based API server using the Axum framework:
Cerebrium supports non-Python apps as long as you provide a Dockerfile. The following example shows a Rust-based API server using the Axum framework:

```rust
use axum::{
Expand Down
2 changes: 1 addition & 1 deletion container-images/custom-web-servers.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ title: "Custom Python Web Servers"
description: Run FastAPI and other ASGI or WSGI Python web servers on Cerebrium with a custom runtime by setting the entrypoint, port, and health check endpoints.
---

Cerebrium's default runtime covers most app needs. For more control, use ASGI or WSGI servers through the custom runtime feature - enabling custom authentication, dynamic batching, frontend dashboards, public endpoints, and WebSocket connections.
Cerebrium's default runtime covers most app needs. For more control, use ASGI or WSGI servers through the custom runtime feature. This enables custom authentication, dynamic batching, frontend dashboards, public endpoints, and WebSocket connections.

## Setting Up Custom Servers

Expand Down
24 changes: 12 additions & 12 deletions container-images/defining-container-images.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ description: Define your Cerebrium container image in cerebrium.toml, from Pytho

## Introduction

Cerebrium abstracts infrastructure management into configuration, so teams focus on app code. A single TOML file manages environment setup, deployments, and scaling — tasks that typically require dedicated teams.
Cerebrium abstracts infrastructure management into configuration, so teams focus on app code. A single TOML file manages environment setup, deployments, and scaling: tasks that typically require dedicated teams.

Unlike traditional Docker or Kubernetes setups with multiple configuration files and orchestration rules, Cerebrium uses a single `cerebrium.toml` file. The system handles container lifecycle, networking, and scaling automatically based on this configuration.

Expand All @@ -18,11 +18,11 @@ Python decorators scatter infrastructure settings throughout code files, making
Run `cerebrium init` to create a `cerebrium.toml` file in the project root. Edit it to match the app's requirements.

<Info>
It is possible to initialize an existing project by adding a `cerebrium.toml`
file to the root of your codebase, defining your entrypoint (`main.py` if
using the default runtime, or adding an entrypoint to the .toml file if using
a custom runtime) and including the necessary files in the `deployment`
section of your `cerebrium.toml` file.
You can initialize an existing project by adding a `cerebrium.toml` file to
the root of your codebase. Define your entrypoint (`main.py` if using the
default runtime, or add an entrypoint to the .toml file if using a custom
runtime) and include the necessary files in the `deployment` section of your
`cerebrium.toml` file.
</Info>

## Hardware Configuration
Expand All @@ -39,7 +39,7 @@ gpu_count = 1 # Number of GPUs

For detailed hardware specifications see the [toml reference](/toml-reference/toml-reference#hardware-configuration).

## Dependency management
## Dependency Management

### Selecting a Python Version

Expand Down Expand Up @@ -86,7 +86,7 @@ Cerebrium caches pip packages at the node level - including wheel files and comp

### Adding APT Packages

System-level packages (image-processing libraries, audio codecs, etc.) are declared under `[cerebrium.dependencies.apt]`:
Declare system-level packages (image-processing libraries, audio codecs, etc.) under `[cerebrium.dependencies.apt]`:

```toml
[cerebrium.dependencies.apt]
Expand Down Expand Up @@ -158,7 +158,7 @@ shell_commands = [
]
```

Use shell commands for tasks that require the fully configured environment — such as compiling code that depends on installed libraries or downloading resources.
Use shell commands for tasks that require the fully configured environment, such as compiling code that depends on installed libraries or downloading resources.

## Custom Docker Base Images

Expand Down Expand Up @@ -271,14 +271,14 @@ vllm = "latest"

### Important Notes

- Code is mounted in `/cortex` - adjust paths accordingly.
- Code is mounted in `/cortex`. Adjust paths accordingly.
- The port in your entrypoint must match the `port` parameter.
- Install any required server packages (uvicorn, gunicorn, etc.) via pip dependencies.
- All endpoints will be available at `https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/your/endpoint`.

Deploy with `cerebrium deploy -y` - the system automatically detects custom runtime configuration.
Deploy with `cerebrium deploy -y`. The system automatically detects custom runtime configuration.

## Deployment process
## Deployment Process

![Deployment process](/images/deployment-process.png)

Expand Down
4 changes: 2 additions & 2 deletions deployments/ci-cd.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ This guide sets up a CI/CD pipeline using GitHub Actions and Cerebrium's Service
<Note>
{" "}
Maintaining **separate development and production apps** in separate projects
is recommended, so that changes can be safely tested before going live.
is recommended, so that you can safely test changes before going live.
</Note>

### 1. Authenticating to Cerebrium
Expand All @@ -23,7 +23,7 @@ The GitHub Action workflow uses a **Service Account key** to authenticate to Cer

![Cerebrium API Keys dashboard](/images/api-keys-sa.png)

### 2. Define secrets in a GitHub environment
### 2. Define Secrets in a GitHub Environment

Store this key in a secret for use in GitHub Actions workflows.

Expand Down
10 changes: 5 additions & 5 deletions deployments/multi-region-deployment.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,10 @@ description: Run a Cerebrium app globally across multiple regions for more GPU c
Deploy an app once and run it in multiple regions. The `region` parameter in `cerebrium.toml` controls placement: run globally on whatever capacity is available (recommended), or pin the app to a specific region. The parameter is optional; when omitted, the platform chooses placement automatically based on the app's hardware requirements.

<Warning>
Multi-region deployment is currently in **beta**. Rapid updates and
improvements will be made over the next few months to bring full functionality
to life. Please reach out on our [Discord](https://discord.gg/ATj6USmeE2)
about features/functionality you would like to see.
Multi-region deployment is currently in **beta**. We will make rapid updates
and improvements over the next few months to bring full functionality to life.
Please reach out on our [Discord](https://discord.gg/ATj6USmeE2) about
features/functionality you would like to see.
</Warning>

## Why Use Multi-Region Deployment
Expand Down Expand Up @@ -125,7 +125,7 @@ The output includes the region of each container.

## Storage

Persistent storage is managed per region: an app has an independent `/persistent-storage` volume in each region it runs in, however placement is configured. Files written in one region are not guaranteed to be available in other regions, and region-local caches, such as model weights downloaded on first load, fill independently per region.
Persistent storage is managed per region: an app has an independent `/persistent-storage` volume in each region it runs in, however placement is configured. Files written in one region are not guaranteed to be available in other regions. Region-local caches, such as model weights downloaded on first load, fill independently per region.

Apps deployed with `region = "global"` also mount `/global-persistent-storage`, a single volume shared across every region the app runs in. Files written there are visible from all regions, and reads are cached per region. Use the global volume for data that must be available everywhere and `/persistent-storage` for region-local data. Manage files on the global volume by passing `--region global` to the file commands. See [Managing Files](/storage/managing-files#global-storage).

Expand Down
8 changes: 4 additions & 4 deletions endpoints/async.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ title: "Async requests"
description: Run Cerebrium functions asynchronously with the async query parameter, get a run_id back instantly, and forward results via a webhook endpoint.
---

Some apps require asynchronous "fire-and-forget" execution. In this model, Cerebrium handles running the function, while the developer is responsible for ensuring data leaves the function (e.g. via a webhook).
Some apps require asynchronous "fire-and-forget" execution. In this model, Cerebrium handles running the function, while you are responsible for ensuring data leaves the function (e.g. via a webhook).

Enable async execution by adding the `async=true` query parameter to the request:

Expand All @@ -30,9 +30,9 @@ X-Request-Id: 21eb3b98-4b10-9ad6-8681-a47172828024
{"run_id":"21eb3b98-4b10-9ad6-8681-a47172828024"}
```

Async functions run for a maximum of **12 hours**, bounded by the `response_grace_period` in `cerebrium.toml`. This defaults to 15 minutes — update it to match the maximum time the task needs.
Async functions run for a maximum of **12 hours**, bounded by the `response_grace_period` in `cerebrium.toml`. This defaults to 15 minutes. Update it to match the maximum time the task needs.

Cerebrium runs the HTTP request in the background, but the function itself must still behave **synchronously** — it must complete its work and return a result.
Cerebrium runs the HTTP request in the background, but the function itself must still behave **synchronously**. It must complete its work and return a result.
Returning a response while the application is still processing causes Cerebrium to begin terminating the container. Only return once all processing is finished.

Because async calls do not return a response to the caller, the function must export any relevant data itself. Combine async execution with a `webhookEndpoint` to have Cerebrium automatically forward the function's response body:
Expand All @@ -44,4 +44,4 @@ curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/<YOUR-APP>/run?async=true&w
--data '{"param": "hello world"}'
```

This is a proxy-level feature — no code changes are required to use webhook forwarding. In the dashboard, the function is marked **async** but still shows the status of the internal synchronous call (e.g. if the call failed, the async request state is `failure`).
This is a proxy-level feature. No code changes are required to use webhook forwarding. In the dashboard, the function is marked **async** but still shows the status of the internal synchronous call (e.g. if the call failed, the async request state is `failure`).
4 changes: 2 additions & 2 deletions endpoints/inference-api.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -3,11 +3,11 @@ title: "REST API"
description: Call your Cerebrium apps over the REST API with POST requests and JWT authentication, and understand response formats and HTTP status codes.
---

All functions on Cerebrium are accessible via POST requests, unless marked private by prefixing the function name with an underscore (e.g. `_private_function()`). Authenticate using the JWT token from the **API Keys** section of the dashboard. Endpoints require this token only when `cerebrium.toml` sets [`disable_auth = false`](/toml-reference/toml-reference) — authentication is disabled by default.
All functions on Cerebrium are accessible via POST requests, unless marked private by prefixing the function name with an underscore (e.g. `_private_function()`). Authenticate using the JWT token from the **API Keys** section of the dashboard. Endpoints require this token only when `cerebrium.toml` sets [`disable_auth = false`](/toml-reference/toml-reference). Authentication is disabled by default.

## Request format

The POST request follows the structure below, where `{function}` is the name of the function to invoke. In this example, `predict()` from `main.py` is called.
The POST request follows the structure below, where `{function}` is the name of the function to invoke. This example calls `predict()` from `main.py`.

```bash
curl --location --request POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/{function}' \
Expand Down
2 changes: 1 addition & 1 deletion endpoints/streaming.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ Streaming sends live output from a model over a server-sent event (SSE) stream.
It works with any Python object that implements the iterator or generator protocol.

The generator/iterator must `yield` data, which is sent downstream via the `text/event-stream` Content-Type.
Data can be sent in JSON format and decoded on the client side.
You can send data in JSON format and decode it on the client side.

A minimal example:

Expand Down
6 changes: 3 additions & 3 deletions endpoints/webhook.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,11 @@ curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/<YOUR-APP>/run?webhookEndpo
--data '{"param": "hello world"}'
```

The proxy forwards the response body as a POST request to the specified webhook — no code changes required. Ensure the webhook endpoint accepts POST requests. Webhook forwarding works with both Cortex and Custom runtimes, but only for HTTP requests (not WebSockets).
The proxy forwards the response body as a POST request to the specified webhook. No code changes are required. Ensure the webhook endpoint accepts POST requests. Webhook forwarding works with both Cortex and Custom runtimes, but only for HTTP requests (not WebSockets).

## Retry Behavior

Webhook requests are sent asynchronously and do not block the function's response. Failed deliveries are retried automatically:
Webhook requests are sent asynchronously and do not block the function's response. Cerebrium retries failed deliveries automatically:

- **Maximum attempts**: 3 attempts total
- **Delay between retries**: Up to 5 seconds between attempts
Expand All @@ -33,7 +33,7 @@ If all attempts fail, the error is logged but will not affect your function's re

<Warning>
**The webhook endpoint should be idempotent.** A webhook may be retried even
if it was already processed — for example, if the Cerebrium backend does not
if it was already processed. For example, the Cerebrium backend may not
receive the success response due to network issues. Design the endpoint to
handle duplicate deliveries gracefully.
</Warning>
Expand Down
4 changes: 2 additions & 2 deletions endpoints/websockets.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ Test the WebSocket endpoint using websocat, a command-line WebSocket client:
websocat wss://api.cerebrium.ai/v4/p-xxxxxxxx/<your-app-name>/<your-websocket-function-name>
```

## Implementing the WebSocket Endpoint
## Implementing the WebSocket endpoint

Example WebSocket endpoint using FastAPI:

Expand All @@ -55,7 +55,7 @@ async def websocket_endpoint(websocket: WebSocket):
await websocket.close()
```

## Additional Info
## Additional info

Client-side Implementation: Handle the WebSocket connection properly on the client, including error handling and reconnection logic.

Expand Down
6 changes: 3 additions & 3 deletions hardware/cpu-and-memory.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -32,8 +32,8 @@ memory = 16.0 # Memory in GB
Allocate system memory equal to the GPU's VRAM capacity as a baseline. This accounts for initial model loading and compilation before GPU transfer. Applications terminate with an Out of Memory (OOM) error if they exceed the specified memory limit.

<Info>
Memory and CPU are billed based on usage, which reduces costs for end-users
and doesn’t require the overprovisioning of an entire instance.
Memory and CPU are billed based on usage, which reduces your costs and doesn’t
require the overprovisioning of an entire instance.
</Info>

## Resource Limits
Expand All @@ -60,4 +60,4 @@ The Transformers library provides memory optimization through the `low_cpu_mem_u

## Resource Monitoring

The platform monitors CPU utilization and throttling events to identify performance bottlenecks. Memory usage and OOM events are tracked to prevent application failures.
The platform monitors CPU utilization and throttling events to identify performance bottlenecks. The platform tracks memory usage and OOM events to prevent application failures.
2 changes: 1 addition & 1 deletion hardware/using-cuda.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ docker_base_image_url = "nvidia/cuda:12.1.1-runtime-ubuntu22.04"

## Cold-Start Optimization

Image size and complexity directly impact cold-start performance — the time needed to initialize an app from an inactive state. Cerebrium uses a content-addressable file system that selectively pulls only required files, but larger images still affect startup times.
Image size and complexity directly impact cold-start performance: the time needed to initialize an app from an inactive state. Cerebrium uses a content-addressable file system that selectively pulls only required files, but larger images still affect startup times.

### Image Size Considerations

Expand Down
4 changes: 2 additions & 2 deletions hardware/using-gpus.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ The platform offers GPUs ranging from cost-effective development options to high
model generation and model name to avoid ambiguity.
</Info>

### Plan availability
### Plan Availability

Compute types are gated by plan:

Expand All @@ -53,7 +53,7 @@ Deploying with a compute type outside the project's plan is rejected at deploy t

## Multi-GPU Configuration

Multiple GPUs are configured in the `cerebrium.toml` file:
Configure multiple GPUs in the `cerebrium.toml` file:

```toml
[cerebrium.hardware]
Expand Down
10 changes: 5 additions & 5 deletions networking/custom-domains.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ Once configured, API calls use the custom domain while keeping the same path str

- Support for apex domains (`example.com`) and subdomains (`api.example.com`)
- Automatic SSL certificate provisioning and renewal
- Project-level domains - one domain serves all apps in a project
- Project-level domains: one domain serves all apps in a project
- Multiple domains can point to the same project
- Professional branding with custom domains instead of `*.cerebrium.ai` URLs

Expand All @@ -32,7 +32,7 @@ Once configured, API calls use the custom domain while keeping the same path str

### Step 2: Configure DNS Records

After creating the domain, DNS configuration instructions will be displayed. Create a CNAME record at the DNS provider using the DNS record Cerebrium generates.
After you create the domain, the dashboard displays DNS configuration instructions. Create a CNAME record at the DNS provider using the DNS record Cerebrium generates.

<Note>
DNS record details can also be found later by clicking "DNS Record" in the
Expand Down Expand Up @@ -78,7 +78,7 @@ After creating the domain, DNS configuration instructions will be displayed. Cre
2. Cerebrium will automatically attempt to validate DNS records every 30 minutes for up to 2 days
3. To trigger an immediate validation attempt, click "Validation Status" -> "Validate Domain" (this works even after the 2-day window has elapsed)
4. If validation fails, the dialog will show the last known error
5. Once validated, SSL certificates will be automatically provisioned
5. Once validated, Cerebrium automatically provisions SSL certificates

### Step 4: Start Using the Custom Domain

Expand Down Expand Up @@ -110,7 +110,7 @@ Cerebrium attempts to validate domains once every 30 minutes for up to 2 days. I

- **Pending**: Domain is waiting for DNS validation
- **Validated**: Domain is successfully validated and ready to use
- **Failed**: DNS validation failed - click "Validation Status" to see the specific error
- **Failed**: DNS validation failed. Click "Validation Status" to see the specific error

### Validation Errors

Expand All @@ -124,7 +124,7 @@ If a domain shows "Failed" status, the DNS record has been misconfigured. Common
- **Cloudflare**: Disable proxy (set to "DNS only" - gray cloud icon)
- **Route53**: Use simple routing policy, not weighted or latency-based
- **Namecheap**: Use "@" for apex domains, not "www" or blank
- **GoDaddy**: CNAME records cannot be used with apex domains - consider using a subdomain or an ALIAS record
- **GoDaddy**: CNAME records cannot be used with apex domains. Consider using a subdomain or an ALIAS record

### Common DNS Mistakes

Expand Down
Loading