Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion calculating-cost.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Calculating compute cost
description: How to calculate the cost of your deployment on Cerebrium
description: Understand how Cerebrium bills GPU, CPU, and memory per second, what counts toward build and runtime charges, and estimate monthly deployment costs.
---

Deployment cost is based on the hardware selected and the execution time. <b>Every time code runs or a machine is specified to stay running, compute is billed</b>. GPU, CPU, and Memory usage are charged per second; persistent storage is charged per GB per month. View compute pricing on the [pricing page](https://www.cerebrium.ai/pricing).
Expand Down
2 changes: 1 addition & 1 deletion container-images/custom-dockerfiles.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Custom Dockerfiles"
description: "Run generic containerized applications on Cerebrium using your own custom Dockerfiles."
description: Deploy containerized apps on Cerebrium with your own Dockerfile, from Python FastAPI servers to compiled Rust binaries, using the custom runtime config.
---

Cerebrium supports deploying existing containerized apps — from standard Python apps to compiled Rust binaries - using a custom Dockerfile. This allows portable, locally reproducible deployment environments.
Expand Down
2 changes: 1 addition & 1 deletion container-images/custom-web-servers.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Custom Python Web Servers"
description: "Run ASGI/WSGI Python apps on Cerebrium"
description: Run FastAPI and other ASGI or WSGI Python web servers on Cerebrium with a custom runtime by setting the entrypoint, port, and health check endpoints.
---

Cerebrium's default runtime covers most app needs. For more control, use ASGI or WSGI servers through the custom runtime feature - enabling custom authentication, dynamic batching, frontend dashboards, public endpoints, and WebSocket connections.
Expand Down
1 change: 1 addition & 0 deletions container-images/defining-container-images.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
---
title: Defining Container Images
description: Define your Cerebrium container image in cerebrium.toml, from Python, pip, apt, and conda dependencies to custom Docker base images and build commands.
---

## Introduction
Expand Down
2 changes: 1 addition & 1 deletion container-images/private-docker-registry.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Using Private Docker Registries"
description: "How to authenticate, pull, and use private Docker images as base images in your deployments."
description: Authenticate with Docker Hub, AWS ECR, or other private registries and use private Docker images as base images for your Cerebrium app deployments.
---

Cerebrium supports private Docker images as base images for deployments, including images from Docker Hub and AWS ECR.
Expand Down
2 changes: 1 addition & 1 deletion deployments/ci-cd.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "CI/CD Pipelines"
description: "Automate Cerebrium deployments using GitHub Actions"
description: Set up a CI/CD pipeline with GitHub Actions and Cerebrium service account keys to automatically deploy your app when a branch is pushed or merged.
---

Configure a Continuous Integration or Continuous Deployment system (CI/CD) to automatically deploy a new version of an app to production/development when a branch is pushed or workflow triggered.
Expand Down
2 changes: 1 addition & 1 deletion deployments/gradual-roll-out.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Gradual Roll-out"
description: "Control the transition between revisions during deployments"
description: Use roll_out_duration_seconds in cerebrium.toml to gradually shift traffic between revisions after a deploy and minimize disruption in production.
---

<Note>This feature is available from CLI version 1.38.2</Note>
Expand Down
2 changes: 1 addition & 1 deletion deployments/multi-region-deployment.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Multi-Region Deployment
description: Deploy an app once and run it across multiple regions, or pin it to a specific region for data residency
description: Run a Cerebrium app globally across multiple regions for more GPU capacity and lower latency, or pin it to one region for data residency needs.
---

Deploy an app once and run it in multiple regions. The `region` parameter in `cerebrium.toml` controls placement: run globally on whatever capacity is available (recommended), or pin the app to a specific region. The parameter is optional; when omitted, the platform chooses placement automatically based on the app's hardware requirements.
Expand Down
2 changes: 1 addition & 1 deletion endpoints/async.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Async requests"
description: "Execute calls to a Cerebrium app to be run asynchronously"
description: Run Cerebrium functions asynchronously with the async query parameter, get a run_id back instantly, and forward results via a webhook endpoint.
---

Some apps require asynchronous "fire-and-forget" execution. In this model, Cerebrium handles running the function, while the developer is responsible for ensuring data leaves the function (e.g. via a webhook).
Expand Down
2 changes: 1 addition & 1 deletion endpoints/inference-api.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "REST API"
description: "Make authenticated HTTP requests to your Cerebrium endpoints"
description: Call your Cerebrium apps over the REST API with POST requests and JWT authentication, and understand response formats and HTTP status codes.
---

All functions on Cerebrium are accessible via POST requests, unless marked private by prefixing the function name with an underscore (e.g. `_private_function()`). Authenticate using the JWT token from the **API Keys** section of the dashboard. Endpoints require this token only when `cerebrium.toml` sets [`disable_auth = false`](/toml-reference/toml-reference) — authentication is disabled by default.
Expand Down
2 changes: 1 addition & 1 deletion endpoints/openai-compatible-endpoints.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "OpenAI-Compatible Endpoints"
description: ""
description: Build an OpenAI compatible chat completions endpoint on Cerebrium and stream responses to the OpenAI Python client using your JWT as the API key.
---

All Cerebrium endpoints are OpenAI-compatible, supporting both `/chat/completions` and `/embedding`. Below is a basic implementation of a streaming OpenAI-compatible endpoint.
Expand Down
1 change: 1 addition & 0 deletions endpoints/streaming.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
---
title: "Streaming Endpoints"
description: Stream live model output from a Cerebrium endpoint over server-sent events by yielding results from a Python generator or iterator function.
---

Streaming sends live output from a model over a server-sent event (SSE) stream.
Expand Down
2 changes: 1 addition & 1 deletion endpoints/webhook.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Webhook Forwarding"
description: "Forward responses to a specified webhook"
description: Forward function responses to an external webhook with a query parameter, including automatic retries and HMAC signature verification on delivery.
---

Forward function response data to an external endpoint via POST by adding the `webhookEndpoint` query parameter to any API call:
Expand Down
1 change: 1 addition & 0 deletions endpoints/websockets.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
---
title: "WebSocket Endpoints"
description: Create real-time bidirectional WebSocket endpoints on Cerebrium using a custom runtime with FastAPI and connect clients over secure wss URLs.
---

WebSocket endpoints stream responses to the client, enabling real-time, bidirectional communication.
Expand Down
1 change: 1 addition & 0 deletions hardware/cpu-and-memory.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
---
title: CPU and Memory
description: Set vCPU cores and memory for Cerebrium apps in cerebrium.toml, review resource limits per hardware type, and optimize usage based billing and OOM risk.
---

## Overview
Expand Down
1 change: 1 addition & 0 deletions hardware/using-cuda.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
---
title: Using CUDA
description: Enable CUDA on Cerebrium using GPU ready Python packages or NVIDIA base images, and choose runtime over devel images to keep cold starts fast.
---

## Overview
Expand Down
2 changes: 1 addition & 1 deletion hardware/using-gpus.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Using GPUs
description: Configure GPU type, quantity, and preference-ordered fallback lists for Cerebrium apps
description: Choose from NVIDIA GPUs like H100, A100, and L40s on Cerebrium, set GPU type and count in cerebrium.toml, and check plan availability for each type.
---

GPUs accelerate computational workloads through parallel processing. Originally designed for graphics rendering, modern GPUs are essential for AI models, large-scale data processing, and other compute-intensive applications.
Expand Down
2 changes: 1 addition & 1 deletion migrations/mystic.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Migrating from Mystic"
description: "Deploy a Model from Mystic on Cerebrium"
description: Migrate your apps from Mystic to Cerebrium with this step by step guide covering config conversion, code changes, deployment, and inference.
---

## Introduction
Expand Down
2 changes: 1 addition & 1 deletion networking/custom-domains.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Custom Domains"
description: "Connect your own domain to your Cerebrium project"
description: Serve Cerebrium apps from your own domain with CNAME setup, DNS validation, automatic SSL certificates, and fixes for common DNS provider issues.
---

Custom domains serve Cerebrium apps through a custom domain instead of the default `*.cerebrium.ai` URLs.
Expand Down
2 changes: 1 addition & 1 deletion other-topics/request-response-logging.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Request and Response Logging"
description: "Control request and response logs in your Cerebrium apps"
description: Disable request and response logging in the Cortex runtime using secrets to protect sensitive data or reduce overhead in Cerebrium apps.
---

By default, the Cortex runtime logs all requests and responses. These logs appear in the app dashboard and are useful for debugging and monitoring. Disable them when privacy, security, or performance requires it.
Expand Down
2 changes: 1 addition & 1 deletion other-topics/using-secrets.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Using Secrets"
description: "Access third-party platforms using secure credentials encrypted on Cerebrium"
description: Store API keys and credentials as encrypted secrets in Cerebrium, expose them as environment variables, and manage them at project or app level.
---

Secrets store API keys, passwords, and other sensitive information outside of code. Secrets are encrypted at rest (256-bit AES) and decrypted only at runtime.
Expand Down
2 changes: 1 addition & 1 deletion partner-services/deepgram.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Deepgram
description: Deploy Deepgram speech-to-text services on Cerebrium
description: Run self hosted Deepgram speech to text on Cerebrium with model file uploads, engine and api TOML setup, GPU scaling, and low latency voice agents.
---

Cerebrium's partnership with [Deepgram](https://www.deepgram.com/) enables simple deployment of speech-to-text (STT) services with simplified configuration and independent scaling.
Expand Down
4 changes: 2 additions & 2 deletions partner-services/index.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Introduction
description: Deploy specialized services from Cerebrium's partners with simplified configurations
title: Partner Services on Cerebrium with Deepgram and Rime
description: Learn how Cerebrium partner services like Deepgram and Rime offer quick deployment, independent scaling, and lower latency for AI workloads.
---

<Note>Partner Services are available from CLI version 1.39.0 and greater</Note>
Expand Down
2 changes: 1 addition & 1 deletion partner-services/rime.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Rime
description: Deploy Rime text-to-speech services on Cerebrium
description: Deploy Rime text to speech on Cerebrium with a simple TOML runtime config, HTTP and WebSocket endpoints, and scaling tuned for low latency TTS.
---

<Note>
Expand Down
2 changes: 1 addition & 1 deletion scaling/batching-concurrency.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Batching and Concurrency"
description: "Improve throughput and cost performance with batching and concurrency"
description: Tune replica_concurrency and use framework native or custom batching with vLLM or LitServe to boost GPU throughput and cut costs on Cerebrium.
---

## Understanding Concurrency
Expand Down
2 changes: 1 addition & 1 deletion scaling/graceful-termination.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Preemption and Graceful Termination"
description: "Implementing Graceful Termination of Instances by Handling Termination Signals"
description: Handle SIGTERM signals in custom runtimes on Cerebrium to finish in flight requests, avoid 502 errors, and shut down instances gracefully.
---

## Graceful Termination
Expand Down
2 changes: 1 addition & 1 deletion storage/managing-files.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Managing Files"
description: "Store model weights and files on regional and global persistent storage volumes"
description: Manage files on Cerebrium persistent storage with CLI upload, download, and list commands, plus global volumes and resizing options up to 1TB.
---

Cerebrium offers file management through a persistent volume that's available to all apps in a project. This storage mounts at `/persistent-storage` and helps store model weights and files efficiently across deployments.
Expand Down
2 changes: 1 addition & 1 deletion toml-reference/toml-reference.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: TOML Reference
description: Complete reference for all parameters available in Cerebrium's default `cerebrium.toml` configuration file.
description: Reference for every cerebrium.toml parameter covering deployment, hardware, scaling, dependencies, and custom runtime settings on Cerebrium.
---

The configuration is organized into the following main sections:
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/asgi-gradio-interface.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Gradio Chat Interface"
description: "Using FastAPI, Gradio and Cerebrium to deploy an LLM chat interface"
description: Deploy a Gradio chat UI for a Llama LLM on Cerebrium with FastAPI and a custom ASGI runtime, running the frontend on CPU while the model scales on GPU.
---

This tutorial covers creating and deploying a Gradio chat interface connected to a Llama 8B language model using Cerebrium's custom ASGI runtime. The architecture runs the frontend on CPU instances while the model runs separately on GPU instances for optimal resource utilization.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/comfyUI.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "ComfyUI application at Scale"
description: "Deploy a ComfyUI application"
description: Turn ComfyUI stable diffusion workflows into autoscaling API endpoints on Cerebrium, from exporting the workflow JSON to serving generated images.
---

### Introduction
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Deploy a Vision Language Model with SGLang"
description: "Build an intelligent ad analysis system that evaluates advertisements across multiple dimensions"
description: Deploy a vision language model with SGLang on Cerebrium and build an ad analysis system that scores advertisements across multiple criteria.
---

This tutorial deploys a Vision Language Model (VLM) using SGLang on Cerebrium. A VLM combines a large language model (LLM) with a vision encoder, enabling it to understand and process both images and text.
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Deploy Triton Inference server and TensorRT-LLM"
description: "Achieve high throughput with Triton Inference Server and the TensorRT-LLM framework"
description: Serve Llama 3.2 with NVIDIA Triton Inference Server and TensorRT-LLM on Cerebrium for up to 15x higher throughput and much lower inference latency.
---

This tutorial deploys Llama 3.2 3B using TensorRT-LLM's PyTorch backend served through Nvidia Triton Inference Server.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/featured.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Featured Examples"
description: "Explore our collection of implementation examples and tutorials"
description: Browse featured Cerebrium examples and tutorials covering LLM endpoints, voice agents, image generation, integrations and other AI applications.
---

<Tabs>
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/gpt-oss.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Serving GPT-OSS with vLLM"
description: "Deploy OpenAI's Latest Open Source Model"
description: Serve OpenAI GPT-OSS open weight models with vLLM on Cerebrium, covering MoE architecture, MXFP4 quantization and H100 GPU deployment setup.
---

GPT recently released GPT-OSS ([gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) and [gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)) two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.0 license, these models outperform similarly sized open models on reasoning tasks, demonstrate strong tool use capabilities, and are optimized for efficient deployment on consumer hardware.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/high-throughput-embeddings.mdx
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
title: "Deploy a High Throughput Server for Embeddings and Reranking"
sidebarTitle: "High-Throughput Embeddings Server"
description: "Deploy a high-throughput, low-latency REST API for serving text-embeddings, reranking models, clip, clap and colpali"
description: Serve text embeddings, reranking, CLIP and ColPali models through a high throughput REST API on Cerebrium using the open source Infinity framework.
---

This tutorial covers deploying a high-throughput, low-latency REST API for serving text-embeddings, reranking models, clip, clap, and colpali using the open-source framework
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/langchain-langsmith.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Langchain and Langsmith"
description: "Deploy an executive assistant using Langsmith and Langchain"
description: Build an executive assistant agent with LangChain tool calling, monitor it in LangSmith and deploy it on Cerebrium to manage Cal.com bookings.
---

This tutorial builds Cal-vin, an executive assistant that manages calendar appointments (via Cal.com) with employees, customers, partners, and friends. It uses the LangChain SDK for agent creation and the LangSmith platform for monitoring scheduling activities and identifying failure points, deployed on Cerebrium for seamless scaling.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/livekit-outbound-agent.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Outbound Agent with LiveKit"
description: "Create an Outbound AI agent that can transfer calls to real agents"
description: Build an outbound AI voice agent with LiveKit and Twilio SIP trunking on Cerebrium that makes calls and warm transfers callers to human agents
---

Voice agents are transforming business operations by introducing efficiencies and personalization for each customer interaction. While most use cases focus on agents receiving calls, this tutorial covers outbound voice AI agents and the use cases they unlock.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/openai-compatible-endpoint-vllm.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "OpenAI compatible vLLM endpoint"
description: "Create an OpenAI compatible endpoint using the vLLM framework"
description: Deploy an OpenAI compatible endpoint for open source LLMs like Llama 3.1 using vLLM on Cerebrium serverless GPUs with streaming responses
---

This tutorial creates an OpenAI-compatible endpoint that works with any open-source model. Use existing OpenAI code with Cerebrium serverless functions by changing just two lines of code.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/realtime-voice-agents.mdx
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
title: "Real-time Voice Agent"
sidebarTitle: "500ms Low-latency Voice Agent"
description: "Deploy a real-time AI voice agent"
description: Build a low latency voice AI agent with PipeCat, Deepgram and a self hosted vLLM Llama endpoint on Cerebrium that responds in about 500ms
---

This tutorial creates a real-time voice agent that responds to queries via speech in ~500ms. The implementation supports swapping in any Large Language Model (LLM) or Text-to-Speech (TTS) model, making it ideal for voice-based use cases like customer support bots and receptionists.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/sdxl.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Generate Images using SDXL"
description: "Generate high quality images using SDXL with refiner"
description: Deploy the Stability AI SDXL refiner model on Cerebrium serverless GPUs to generate high quality images from prompts with a Diffusers pipeline
---

<Note>
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/transcribe-whisper.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Transcribe 1 hour podcast"
description: "Using Distill Whisper to transcribe an audio file"
description: Transcribe hour long podcasts and audio files with Distil Whisper on Cerebrium using base64 uploads or file URLs and webhooks for long jobs
---

This tutorial transcribes an hour-long audio file using Distill Whisper — an optimized version of Whisper-large-v2 that's 60% faster while maintaining accuracy within 1% of the original. The endpoint accepts either a base64-encoded string of the audio file or a URL to download the audio file.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/twilio-voice-agent.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Twilio Voice Agent with PipeCat"
description: "Integrate a real-time AI voice agent with Twilio"
description: Connect a real time AI voice agent to phone calls using Twilio, PipeCat and FastAPI WebSockets on Cerebrium for support bots and receptionists
---

This tutorial creates a real-time voice agent that responds to phone calls via Twilio. The implementation supports any LLM or Text-to-Speech (TTS) model, making it ideal for voice applications like customer support bots and receptionists.
Expand Down
2 changes: 1 addition & 1 deletion v4/examples/wandb-sweep.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Hyperparameter Sweep training Llama 3.2 with WandB"
description: "Run a hyperparameter sweep on Llama 3.2 with WandB"
description: Fine tune Llama 3.2 with Weights and Biases hyperparameter sweeps, running training experiments in parallel across Cerebrium serverless GPUs
---

Hyperparameter sweeps systematically test parameter combinations to find the best-performing model for the least compute or training time.
Expand Down
Loading