An open-source LLM inference platform for Kubernetes: deploy a model with a tuned configuration, route each request to the replica that already holds its context, scale on LLM signals and SLOs, and tune the serving configuration on your own workload.
Start here: modelsphere/modelsphere — the deployment repository, with the install guide and the architecture.
| Area | Repository | What it does |
|---|---|---|
| Install | modelsphere | Installs the whole stack: Ansible, helmfile, offline bundle |
| helm-charts | Helm charts for the sglang and vllm engines and the routing components | |
| model-catalog | The models swiss can deploy, and how to serve each one: engine, image, flags, GPUs, tuned variants | |
| Routing | llm-openresty | Session-affinity router, request logging and metrics |
| cache_aware_router | Routes to the replica holding the longest matching prompt prefix | |
| autoconfig | Keeps the routing layer's config in sync with the model backends | |
| Scaling and health | llm-operator | Autoscales on KV-cache utilization, queue depth and TPM |
| slo-scaler-decision-gen | Turns SLO targets and live signals into replica decisions | |
| slo-api | HTTP API to read and change a service's SLO | |
| hang-watcher | Restarts an engine that stopped making progress | |
| Operate | swiss | Deploy control plane for the engine charts |
| console | Web console: users and roles, model deployment, playground | |
| Tune and measure | llm-autotune | Searches serving configurations on idle GPUs |
| llm-autotune-policies | Search policies and the policy SDK for llm-autotune | |
| llm-bench | Benchmark platform for LLM serving endpoints |
Everything here is licensed under Apache-2.0. Contributing · Code of Conduct · Security