Conversation
Route tokens locally in expert-major order, combine their outputs without an EpBuffer or NCCL EP, and validate unsupported communication settings. Cover the five-op MoE forward and backward path against a dense reference. Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
for more information, see https://pre-commit.ci
13 tasks
vthumbe1503
pushed a commit
that referenced
this pull request
Sep 29, 2026
…IA#3137) * start draft Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove benchmark scripts Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * add license Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make cutlass dsl required Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix linting errors Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * warn failed compilation Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make tvm-ffi common dependency Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * use a higher version for cutlass-dsl Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * less comment Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * skip zeroing buffer if noop flag set Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * refactor activations Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * lint Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * maybe it's better to prepare tvm-ffi & cutedsl kernels in common/__init__.py Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * also flush colwise scale in SMEM Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * use explicit fp8 types Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * add checks Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * refactor Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fi Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * skip comparision test if cutedsl disabled Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * Merge pull request #3 from janekb04/cutedsl_mxfp8_common Perform noop tensor check on device Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * add test to the qa script Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * warn only once if dlopen fails Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * let tests cover swizzling Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * warn not chosen cutedsl kernel during testing Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * improve lazyload so now the cache is uint32 Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * let tests be verbose about warnings so we know if cutedsl kernel is not dispatched Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * support swizzled scale for specialized kernel but disable it Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix inline ptx Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * more fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * revert to mul Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add jax tests Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * also ignore noop when dbias Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * relax dim check for the kernels Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * reject 2D quant Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * run more tests with CuTeDSL path Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * make rowwise only handle divisible cases Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fallback to CUDA if not TMA aligned Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * specialize skip masking version Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add type annotations Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * optimize masking Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * skip masking for some acts Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * allow user to opt out from building with CuTeDSL Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * better warning Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * doc Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add jax test coverage for cutedsl Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix device sm query Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * catch more python error Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make tvm-ffi required for python users Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * optimize dbias Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * skip scale bound check whenever possible Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * allow colwise to fuse relu too Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * dispatch to rowwise specialized with swizzled scale Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * don't link tvmffi if building a standalong C++ lib without framework Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * change some logging Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * use the right fast math flag Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * also let test script use NVTE_DEBUG Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * also let mutex cover the loading of tvm-ffi func Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * locate tvm-ffi include dir from setup_common_extension Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix how tests run with cutedsl enabled Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * reduce framework scope tests run with cutedsl backend Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * allow C++ tests to run with cutdsl Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove tvmffi and cutlass dependencies from setup.py due to being optional Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * encode sm numbers in tvm_ffi function name & cache key Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make C++ tests run in sequential Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make CuTeDSL default and cover C++ in L1 Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * Improvements and fix to CuTe DSL kernel launches (#7) * [Common] Bypass Python wrapper overhead with __tvm_ffi_object__ Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Avoid heap allocations in DLTensorWrapper Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Avoid querying device in DLTensorWrapper Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Fix packed dtype handling in DLTensorWrapper Signed-off-by: Jan Bielak <jbielak@nvidia.com> * Apply review suggestions Signed-off-by: Jan Bielak <jbielak@nvidia.com> --------- Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix dispatching so it never uses specialized if grid won't fit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * also add tests Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * Minor fixes to CuTe DSL infrastructure (NVIDIA#8) * Improvements and fix to CuTe DSL kernel launches (#7) * [Common] Bypass Python wrapper overhead with __tvm_ffi_object__ Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Avoid heap allocations in DLTensorWrapper Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Avoid querying device in DLTensorWrapper Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Fix packed dtype handling in DLTensorWrapper Signed-off-by: Jan Bielak <jbielak@nvidia.com> * Apply review suggestions Signed-off-by: Jan Bielak <jbielak@nvidia.com> --------- Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * [Common] Add missing dtypes to CuTe DSL dispatch Signed-off-by: Jan Bielak <jbielak@nvidia.com> * [Common] Add line information to CuTe DSL inline PTX helpers Signed-off-by: Jan Bielak <jbielak@nvidia.com> --------- Signed-off-by: Jan Bielak <jbielak@nvidia.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add missing ip & loc Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * refactor so we don't need #pragma to avoid collision in tests Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * whoops Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * whoops again Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * rename NVTE_ENABLE_CUTEDSL_QUANT_BACKEND to NVTE_ENABLE_CUTEDSL_BACKEND Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * switch back to NVTE_WARN_IF_CUTEDSL_BACKEND_NOT_CHOSEN Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * use TE's to_string to convert dtype to str Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove MXFP8QuantFused Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * move python init from tests to TE lib Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * software path for e8m0 conversion Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove unreachable branch Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fold quantize_mxfp8_cutedsl_config.h back to quantize_mxfp8_cutedsl.h Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit for some C++ Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * rewrite dsl_user_op to cute.jit if possible and let compiler optimize Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix python init Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * prefere inline_ptx instead of llvm.inline_asm Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix activation encoding Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove unused device_is_blackwell function Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * link C++ test executable to libtvm_ffi.so Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * move cutedsl utils from math.h to a separate file Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * recover test_cast_mxfp8.cu Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * no need to init extension now Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * hoist importorskip in cutedsl python tests Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix activation_func_to_enum so it works with CUDA 12.8 Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make NVTE_WITH_CUTEDSL both a env var and a cmake option Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix TE path searching for python bootstrap Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * when enabling cutedsl manually, also initialize python if it's not done before Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * use dlsym Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * let main TE lib link to Python::Module Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * add a test to cover untested manually enable cutedsl path Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * update min ver workflow Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * rename L1 CUDA C++ UT Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix build issues for CuTeDSL backend (NVIDIA#10) * disable CuTeDSL support and other fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove docs that mention CuTeDSL since they are off by default now Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * remove test_cutedsl_backend.cpp Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * make envvar.rst more human Signed-off-by: Kaining Zhong <kainingz@nvidia.com> --------- Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix python tests Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * skip L1 C++ UT with CUDA when CuTeDSL is not enabled globally Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix C++ UT python discovery Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix deadlock for multigpu Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix python init Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * fix python init again thanks to codex Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * refactor the serialized initial launch Signed-off-by: Kaining Zhong <kainingz@nvidia.com> * nit Signed-off-by: Kaining Zhong <kainingz@nvidia.com> --------- Signed-off-by: Kaining Zhong <kainingz@nvidia.com> Signed-off-by: Jan Bielak <jbielak@nvidia.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Jan Bielak <janekb04@icloud.com> Co-authored-by: Jan Bielak <jbielak@nvidia.com>
…VIDIA#3467) * [Common] Add single-launch grouped fused amax for row-scaled NVFP4 Compute per-expert rowwise/columnwise amax over a packed (sum_M, K) grouped input in one kernel launch instead of one launch per expert, exposed via nvte_group_nvfp4_compute_amax and gated by NVTE_NVFP4_FUSED_AMAX. Adds a gtest that checks the result against a CPU reference and across launches. Signed-off-by: Cael Ling <caell@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * [Common] Enforce K alignment and per-expert amax buffers in grouped NVFP4 amax The public entry point calls the launcher directly, bypassing the eligibility check, so misaligned K silently dropped trailing columns and heterogeneous per-expert amax buffers could dereference null. Add a column/split alignment check in the launcher, take do_row/do_col as the union over experts with a per-expert null guard in the kernel, and cover it with HeterogeneousColumnwise. Signed-off-by: Cael Ling <caell@nvidia.com> * [Common] nvfp4: drop unused grouped fused-amax gate and env switch group_fused_amax_supported was never called (the C API calls the launcher directly and alignment checks already live there), and it carried the NVTE_NVFP4_FUSED_AMAX switch. Remove the dead function and its now-unused <cstdlib>. Signed-off-by: Cael Ling <caell@nvidia.com> --------- Signed-off-by: Cael Ling <caell@nvidia.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
NVIDIA#3553) * Isolate FlashAttention 2/3 imports so a broken install cannot break TE The version gate only inspects distribution metadata, so a flash-attn whose extension fails to load -- missing CUDA dependency, ABI mismatch -- raised straight out of backends.py and made `import transformer_engine.pytorch` fail outright, even for callers that never ask for FlashAttention. FlashAttention 4 already guards its import this way, so give FA2 and FA3 the same treatment: on ImportError leave the module aliases at None, mark the backend unavailable, and warn with the version and the original error. A genuine TE failure still raises; only the optional dependency is contained. The regression test injects a metadata-only flash-attn distribution in a subprocess, because the fault has to be in place before transformer_engine is imported. It fails for both FA2 and FA3 without this change and passes with it. Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de> * Fix optional FlashAttention import regression coverage Signed-off-by: Przemek Tredak <ptredak@nvidia.com> --------- Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de> Signed-off-by: Przemek Tredak <ptredak@nvidia.com> Co-authored-by: Przemek Tredak <ptredak@nvidia.com>
…A#3565) * [common] Reject mismatched bias/out dtype for non-FP8/FP4 GEMM cuBLASLt assumes bias dtype equals D when BIAS_DATA_TYPE is unset. TE only sets that attribute on the FP8/FP4 path, so BF16 bias with FP32 out was reinterpreted and read past the bias allocation. Match cuBLAS support tables and Megatron-LM#6000: require bias dtype == D on the non-FP8/FP4 path. Fixes NVIDIA#3562 Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * [qa] Wire test_gemm_bias_dtype into L0 pytorch unittest L0 enumerates explicit pytest paths, so the new bias dtype regression module was never invoked by community CI. Add it next to the other top-level pytorch tests. Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com> * [qa] Move non-fp8 bias dtype check into test_sanity Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com> --------- Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Przemyslaw Tredak <ptredak@nvidia.com>
Signed-off-by: Santosh Bhavani <santosh.bhavani@live.com>
) * [PyTorch] MoE Sequential block with Dispatch and Combine as basic ops Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com> Signed-off-by: Varun Thumbe <vthumbe@nvidia.com> * minor bug fix in test Signed-off-by: Varun Thumbe <vthumbe@nvidia.com> --------- Signed-off-by: Varun Thumbe <vthumbe@nvidia.com> Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
…iled kernel (NVIDIA#3454) * [Common] nvfp4: fuse row/col amax into a single TMA-tiled kernel The row-scaled path ran two amax kernels; the columnwise one read global memory column-major (uncoalesced) and dominated runtime. Compute both directions in one kernel that streams 128x128 chunks through shared memory via TMA and reduces columns from SMEM. Gated by fused_amax_supported (BF16, 128-aligned dims); other cases keep the two-kernel path. Set NVTE_NVFP4_FUSED_AMAX=0 to force fallback. Signed-off-by: Cael Ling <caell@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * [Common] Preserve NVFP4 fused amax buffers on no-op quantize The fused wrapper zeroed both amax buffers with an unconditional memset while the kernel returns early on noop[0]==1, so a skipped (graph-replay) call cleared the previously published amax. Replace the memset with a noop-aware zero kernel that returns early on the same flag, matching the standalone kernels' contract. Signed-off-by: Cael Ling <caell@nvidia.com> * [Common] nvfp4: hide row-scaled amax fused/unfused behind one entry Drop the NVTE_NVFP4_FUSED_AMAX switch and make fused-vs-unfused an internal detail; callers use nvfp4::row_scaled::compute_amaxes, unfused stays as fallback. Signed-off-by: Cael Ling <caell@nvidia.com> * Assume fused amax kernel computes both row-wise and col-wise amaxes Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com> Signed-off-by: Tim Moon <4406448+timmoon10@users.noreply.github.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Signed-off-by: Cael Ling <caell@nvidia.com> Signed-off-by: Tim Moon <4406448+timmoon10@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add a bufferless PyTorch backend for
MoeDispatchandMoeCombinewhen the expert-parallel group has one rank. Dispatch orders local routes by expert and combine sums expert outputs back into source tokens. Routing metadata belongs to each invocation, so two outstanding forwards do not overwrite each other. The local backend rejects remote expert IDs and communication settings it cannot honor.This is the EP=1 follow-up requested on NVIDIA/TransformerEngine#3503. This draft is stacked on
vthumbe1503:nccl_ep_opsat261a81c; its signed-off implementation commit and pre-commit.ci formatting commit change only the EP=1 path, tests, and API documentation. The main-targeting upstream draft is open.Validation
CUDA_VISIBLE_DEVICES=3 PYTHONPATH=$PWD python -m pytest tests/pytorch/test_moe_ep_local.py -q6ce761bMoeDispatch → GroupedLinear → ScaledSwiGLU → GroupedLinear → MoeCombinertol=atol=2e-2, including token, gate-weight, and expert-weight gradientsThe test imported the PR source tree, ran without an initialized
torch.distributedprocess group, and the loaded native extension exposed no EP entry points. The NCCL EP distributed suite was not run on this machine. The full L0 launcher was not run because it installs dependencies; the focused tests used the existing environment without package changes.