Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches
-
Updated
Jul 12, 2026 - Python
Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches
Qwen3.8-Flash-Next on 4x RTX 3090 with vLLM: 806,792-token KV pool, three 262K sessions resident, MTP, host-mapped PLE, pinned build and container recipes.
Qwen3.8 Flash Next FP8 on 4x CMP 170HX: PP3 transformer stages with the 51B PLE table on GPU 3
A ~800-line PyTorch implementation of Megatron-LM's TP + PP + DP + AMP. 1.6-2x faster than Megatron-Core on 125M models.
FSDP · DDP · Pipeline · Tensor parallel on 4× NVIDIA A30 — real throughput benchmarks for Qwen2.5-7B multi-GPU training
Xe2 dual-B70 kernel and 2x2 parallelism lab (TP=2, PP=2)
JAX/Flax NNX training infrastructure: device meshes and SPMD sharding, an Orbax checkpoint store, early stopping and callbacks, W&B and MLflow logging.
To associate your repository with the pipeline-parallel topic, visit your repo's landing page and select "manage topics."