[ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
-
Updated
May 1, 2025 - Python
[ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU (DAC ‘26)
CPU-side action routing, context compression, and causal memory for AI agents — matches an LLM-everything agent at 58% fewer LLM calls and ~45% lower cost. Glues busyBee-cpu, honey-comb, and rust-brain.
Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
DACAN (Дацан): Qwen3.8-Flash-Next (125B MoE) on any NVIDIA RTX 20/30/40/50 card and any x86-64 CPU with AVX2, AVX-512 not required. One or two cards, one or two sockets, context up to 524K.
Run a closure in a child process and return the result over a promise
Run closures in a child process messenger or pool
[arXiv] ColA: Collaborative Adaptation with Gradient Learning
QSA HiSparse for SGLang: 256K KV offload, CUDA Graph benchmarks, and patches tested on dual RTX 4090 48GB
Measurement-first research on local Mixture-of-Experts inference under a hardware contract. 3 measured laws, 4 falsified ideas, and paper site.
Engine for two MoE models only - DeepSeek-V4.1-Flash (main) and GLM-5.3-Flash - on one 96 GB GPU, experts offloaded to CPU RAM, 262K context. Pieced together from what we had (DDR4, PCIe 4); DDR5 would do better. Sleeps/wakes in seconds to share the GPU. OpenAI/Anthropic API, works behind LiteLLM.
Run DeepSeek-V4.1-Flash on 8x RTX 5090 + 503 GiB RAM. 5.6x output throughput vs patched eager mode, with vLLM fixes, CPU offload & reproducible benchmarks.
Recall-aware tiered KV cache engine for long-context local LLM inference.
DeepSeek-V4.1-Flash at TP3 on 3x RTX PRO 6000 (SM120): EXL3 vs official checkpoint + UVA offload, same-box benchmarks, failed attempts, reproducible configs.
To associate your repository with the cpu-offload topic, visit your repo's landing page and select "manage topics."