Hands-on labs in LLM inference: serving, caching, quantization, throughput. Every lab runs on free-tier GPUs, ships a notebook, and reports measured numbers — no vibes, only benchmarks.
| # | Lab | Result |
|---|---|---|
| 01 | Prefix caching on vLLM (Qwen3-4B, T4) | 7.54× prefill speedup from KV-block reuse |
| 02 | Quantization: AWQ-4bit vs fp16 (Qwen3-4B, T4) | 3× smaller, 2.25× KV room, 0× faster |
| 03 | Throughput sweep: ceiling, knee, cliff (Qwen3-4B-AWQ, T4) | ~3 req/s sellable @ 128 ms TTFT; 510 tok/s only @ 3.6 s p99 |
| 04 | Chunked prefill: killing the ITL freeze (Qwen3-4B-AWQ, T4) | 31× reduction in freeze (2.7s to 86ms), 4× drop in P99 ITL |
| 06 | KV cache behavior, capacity & preemption (Qwen3-4B-AWQ, T4) | 34.6k token ceiling verified, 98% APC masking, graceful degradation vs OOM |
Lab 05 (SGLang and RadixAttention vs vLLM prefix caching) is next.
Each lab is also written up as a short post on LinkedIn.
Each lab follows the same contract:
- Serve a real model on a real (free) GPU
- Benchmark with confounds removed (warmup, fixed seeds, isolated variables)
- Report cold vs warm / before vs after with identical outputs
- State exactly how to reproduce for $0