Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
-
Updated
Sep 30, 2026 - Swift
Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.
Research preview for reproducible MoE routing and memory oversubscription on consumer GPUs. Not a production inference engine.
Flash-backed mixture-of-experts inference on Apple devices with Swift/MLX.
DualDeadline adds separate gate/up and down-projection transfer deadlines to exact MoE offloading, with H200 validation and Triton-optimized predictors.
Stream PyTorch model blocks from NVMe or pinned CPU memory with a bounded GPU residency budget. LoRA-finetune models far larger than host RAM on one GPU - Qwen3-VL-32B and gpt-oss-120b train in under 6 GB of VRAM.
Research artifact for Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency
Run BF16 LLMs larger than your GPU's memory without quantization: lossless weight compression, layer streaming, and speculative decoding. Code and frozen evidence for H6.5, schedule-aware exact weight placement.
To associate your repository with the model-offloading topic, visit your repo's landing page and select "manage topics."