NVIDIA Tesla V100 显卡驱动自动化管理套件
-
Updated
Oct 9, 2025 - PowerShell
NVIDIA Tesla V100 显卡驱动自动化管理套件
PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16GB card. Fork of ik_llama.cpp.
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
Open hardware desktop AI node: 4× Tesla V100, 128GB HBM2, PCIe/NVLink topology and V-Core liquid/air cooling.
Run Mistral Voxtral realtime speech-to-text on a Tesla V100 (sm_70/Volta) via a patched vLLM — RTF ~0.4, keeps up real time
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
Add a description, image, and links to the tesla-v100 topic page so that developers can more easily learn about it.
To associate your repository with the tesla-v100 topic, visit your repo's landing page and select "manage topics."