llm-d
Here are 28 public repositories matching this topic...
Self-hosted, distributed, multi-model LLM inference
-
Updated
Sep 21, 2026 - Go
Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping
-
Updated
Sep 24, 2026 - Go
Distributed inference configuration tools
-
Updated
Sep 24, 2026 - TypeScript
Accelerate reproducible inference experiments for large language models with LLM-D! This lab automates the setup of a complete evaluation environment on OpenShift/OKD: GPU worker pools, core operators, observability, traffic control, and ready-to-run example workloads.
-
Updated
Sep 1, 2026 - Jupyter Notebook
Experimental recovery coordination for expert-parallel inference groups in llm-d.
-
Updated
Sep 10, 2026 - Go
Stack to serve LLM inference based on llm-d, Envoy AI gateway
-
Updated
Sep 22, 2026 - Go
Prove which open model can replace your closed one — on your traffic, your hardware, your budget. Then move you there without a risky cutover.
-
Updated
Aug 18, 2026 - Python
Self-hosted LLM inference platform on AWS EKS — vLLM, llm-d, Cosign, ArgoCD, Kyverno, Kagent+MCP
-
Updated
Sep 9, 2026 - HCL
RecursiveCharacterTextSplitter and context cache with llm-d
-
Updated
Jan 29, 2026 - Python
Open Source software code for use with PCIe card-based hardware AI accelerators catering to both inference and training use cases
-
Updated
Sep 12, 2026 - Python
LLM inference stack on Kubernetes: Envoy + EPP router + vLLM on minikube, CPU-only — benchmarked against a naive HF server
-
Updated
Aug 27, 2026 - Python
Verifiable, tenant-isolated LLM inference controls for regulated AI infrastructure.
-
Updated
Sep 13, 2026 - Python
Add this topic to your repo
To associate your repository with the llm-d topic, visit your repo's landing page and select "manage topics."