Distill a repeated Jev classification task into a local model, on the fly — same answers, your hardware.
-
Updated
Oct 1, 2026 - Python
Distill a repeated Jev classification task into a local model, on the fly — same answers, your hardware.
Distil an expensive LLM API call on a narrow task into a small local model: capture traffic, curate, LoRA fine-tune with MLX on Apple Silicon, evaluate against the teacher, and serve an OpenAI-compatible cascade that escalates low-confidence requests.
Cost-aware routing across agentic pipelines. Empirical benchmarks of accuracy, cost, and dangerous rate.
Studying Minimum Sufficient Inference: when objective execution evidence can stop LLM inference without sacrificing reliability.
Adaptive Cascade Tuning via Importance Sampling
NSGA-II search framework for CIFAR-10 big/little dynamic inference cascades under embedded memory constraints.
LLM serving router that prices every request in dollars and milliseconds. Semantic cache plus a confidence-gated cheap-to-GPT-4 cascade: 96.9% lower cost than always calling the strong model, within 1.6 accuracy points, median latency 450ms to sub-ms on repeated traffic. FastAPI, SQLite cost ledger, reproducible benchmark.
Ask the cheap model first. Pay for the strong one only when the cheap answer fails a check, under a budget, with a cost record for every call.
Route each LLM call to the cheapest model that will get it right. Verifier-gated cascades with an offline threshold optimizer — 94% of the strongest model's accuracy at 34% lower cost.
Divergence-aware multi-agent routing that cuts LLM inference cost 10-100x. One signal routes queries, keys a multilingual cache, and detects hallucinations at AUC 0.90. CIKM 2026.
To associate your repository with the model-cascade topic, visit your repo's landing page and select "manage topics."