Skip to content
#

cpu-offload

Here are 17 public repositories matching this topic...

Engine for two MoE models only - DeepSeek-V4.1-Flash (main) and GLM-5.3-Flash - on one 96 GB GPU, experts offloaded to CPU RAM, 262K context. Pieced together from what we had (DDR4, PCIe 4); DDR5 would do better. Sleeps/wakes in seconds to share the GPU. OpenAI/Anthropic API, works behind LiteLLM.

  • Updated Oct 7, 2026
  • C++

Add this topic to your repo

To associate your repository with the cpu-offload topic, visit your repo's landing page and select "manage topics."

Learn more