MoE expert offload for low-VRAM GPUs — run 100B+ MoE models (DeepSeek, Qwen, Mixtral) on 8 GB cards. Expert cache with LRU hot experts, OpenAI-compatible proxy, GGUF multi-shard.
-
Updated
Sep 12, 2026 - Python
MoE expert offload for low-VRAM GPUs — run 100B+ MoE models (DeepSeek, Qwen, Mixtral) on 8 GB cards. Expert cache with LRU hot experts, OpenAI-compatible proxy, GGUF multi-shard.
Run AI models too large for your Mac's memory — at near-full speed. Intelligent expert caching, speculative execution, and 15+ research techniques for MoE inference on Apple Silicon.
Flash-backed mixture-of-experts inference on Apple devices with Swift/MLX.
Run large MLX models on Apple Silicon with flash weight streaming, using native precision beyond RAM limits
To associate your repository with the expert-caching topic, visit your repo's landing page and select "manage topics."