Skip to content

Add a PyTorch EP=1 backend for MoE dispatch and combine - #7

Draft
0z5a wants to merge 10 commits into
vthumbe1503:nccl_ep_opsfrom
0z5a:ep1-pytorch-dispatch-combine
Draft

0z5a wants to merge 10 commits into
vthumbe1503:nccl_ep_opsfrom
0z5a:ep1-pytorch-dispatch-combine

Conversation

@0z5a

@0z5a 0z5a commented Sep 28, 2026 •

Copy link
Copy Markdown

Summary

Add a bufferless PyTorch backend for MoeDispatch and MoeCombine when the expert-parallel group has one rank. Dispatch orders local routes by expert and combine sums expert outputs back into source tokens. Routing metadata belongs to each invocation, so two outstanding forwards do not overwrite each other. The local backend rejects remote expert IDs and communication settings it cannot honor.

This is the EP=1 follow-up requested on NVIDIA/TransformerEngine#3503. This draft is stacked on vthumbe1503:nccl_ep_ops at 261a81c; its signed-off implementation commit and pre-commit.ci formatting commit change only the EP=1 path, tests, and API documentation. The main-targeting upstream draft is open.

Validation

Check Result
CUDA_VISIBLE_DEVICES=3 PYTHONPATH=$PWD python -m pytest tests/pytorch/test_moe_ep_local.py -q 13 passed in 11.57 s on RTX 5090, PyTorch 2.13.0+cu130, at head 6ce761b
Five-op MoeDispatch → GroupedLinear → ScaledSwiGLU → GroupedLinear → MoeCombine BF16 forward and backward match an independent dense PyTorch MoE within the test's rtol=atol=2e-2, including token, gate-weight, and expert-weight gradients
Twelve-step training E2E with changing inputs and top-k=2 routes Full five-op forward, backward, SGD updates, and post-update inference matched the dense PyTorch reference; 8 parameter tensors changed, maximum output absolute difference 0.00178438 and maximum gradient absolute difference 0.00277808
Routing cases top-k=1 and 2, repeated and dropped routes, empty experts, two outstanding forwards, and unsupported settings covered
Repository license checker Passed

The test imported the PR source tree, ran without an initialized torch.distributed process group, and the loaded native extension exposed no EP entry points. The NCCL EP distributed suite was not run on this machine. The full L0 launcher was not run because it installs dependencies; the focused tests used the existing environment without package changes.

Route tokens locally in expert-major order, combine their outputs without an EpBuffer or NCCL EP, and validate unsupported communication settings. Cover the five-op MoE forward and backward path against a dense reference.

Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
vthumbe1503 pushed a commit that referenced this pull request Sep 29, 2026
…IA#3137)

* start draft

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove benchmark scripts

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* add license

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make cutlass dsl required

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix linting errors

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* warn failed compilation

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make tvm-ffi common dependency

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* use a higher version for cutlass-dsl

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* less comment

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* skip zeroing buffer if noop flag set

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* refactor activations

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* lint

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* maybe it's better to prepare tvm-ffi & cutedsl kernels in common/__init__.py

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* also flush colwise scale in SMEM

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* use explicit fp8 types

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* add checks

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* refactor

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fi

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* skip comparision test if cutedsl disabled

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* Merge pull request #3 from janekb04/cutedsl_mxfp8_common

Perform noop tensor check on device

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* add test to the qa script

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* warn only once if dlopen fails

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* let tests cover swizzling

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* warn not chosen cutedsl kernel during testing

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* improve lazyload so now the cache is uint32

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* let tests be verbose about warnings so we know if cutedsl kernel is not dispatched

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* support swizzled scale for specialized kernel but disable it

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix inline ptx

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* more fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* revert to mul

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* add jax tests

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* also ignore noop when dbias

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* relax dim check for the kernels

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* reject 2D quant

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* run more tests with CuTeDSL path

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* make rowwise only handle divisible cases

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fallback to CUDA if not TMA aligned

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* specialize skip masking version

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* add type annotations

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* optimize masking

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* skip masking for some acts

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* allow user to opt out from building with CuTeDSL

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* better warning

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* doc

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* add jax test coverage for cutedsl

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix device sm query

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* catch more python error

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make tvm-ffi required for python users

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* optimize dbias

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* skip scale bound check whenever possible

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* allow colwise to fuse relu too

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* dispatch to rowwise specialized with swizzled scale

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* don't link tvmffi if building a standalong C++ lib without framework

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* change some logging

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* use the right fast math flag

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* also let test script use NVTE_DEBUG

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* also let mutex cover the loading of tvm-ffi func

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* locate tvm-ffi include dir from setup_common_extension

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix how tests run with cutedsl enabled

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* reduce framework scope tests run with cutedsl backend

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* allow C++ tests to run with cutdsl

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove tvmffi and cutlass dependencies from setup.py due to being optional

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* encode sm numbers in tvm_ffi function name & cache key

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make C++ tests run in sequential

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make CuTeDSL default and cover C++ in L1

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* Improvements and fix to CuTe DSL kernel launches (#7)

* [Common] Bypass Python wrapper overhead with __tvm_ffi_object__

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Avoid heap allocations in DLTensorWrapper

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Avoid querying device in DLTensorWrapper

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Fix packed dtype handling in DLTensorWrapper

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* Apply review suggestions

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

---------

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix dispatching so it never uses specialized if grid won't fit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* also add tests

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* Minor fixes to CuTe DSL infrastructure (NVIDIA#8)

* Improvements and fix to CuTe DSL kernel launches (#7)

* [Common] Bypass Python wrapper overhead with __tvm_ffi_object__

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Avoid heap allocations in DLTensorWrapper

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Avoid querying device in DLTensorWrapper

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Fix packed dtype handling in DLTensorWrapper

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* Apply review suggestions

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

---------

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [Common] Add missing dtypes to CuTe DSL dispatch

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

* [Common] Add line information to CuTe DSL inline PTX helpers

Signed-off-by: Jan Bielak <jbielak@nvidia.com>

---------

Signed-off-by: Jan Bielak <jbielak@nvidia.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* add missing ip & loc

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* refactor so we don't need #pragma to avoid collision in tests

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* whoops

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* whoops again

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* rename NVTE_ENABLE_CUTEDSL_QUANT_BACKEND to NVTE_ENABLE_CUTEDSL_BACKEND

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* switch back to NVTE_WARN_IF_CUTEDSL_BACKEND_NOT_CHOSEN

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* use TE's to_string to convert dtype to str

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove MXFP8QuantFused

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* move python init from tests to TE lib

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* software path for e8m0 conversion

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove unreachable branch

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fold quantize_mxfp8_cutedsl_config.h back to quantize_mxfp8_cutedsl.h

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit for some C++

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* rewrite dsl_user_op to cute.jit if possible and let compiler optimize

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix python init

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* prefere inline_ptx instead of llvm.inline_asm

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix activation encoding

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove unused device_is_blackwell function

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* link C++ test executable to libtvm_ffi.so

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* move cutedsl utils from math.h to a separate file

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* recover test_cast_mxfp8.cu

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* no need to init extension now

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* hoist importorskip in cutedsl python tests

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix activation_func_to_enum so it works with CUDA 12.8

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make NVTE_WITH_CUTEDSL both a env var and a cmake option

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix TE path searching for python bootstrap

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* when enabling cutedsl manually, also initialize python if it's not done before

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* use dlsym

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* let main TE lib link to Python::Module

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* add a test to cover untested manually enable cutedsl path

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* update min ver workflow

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* rename L1 CUDA C++ UT

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix build issues for CuTeDSL backend (NVIDIA#10)

* disable CuTeDSL support and other fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove docs that mention CuTeDSL since they are off by default now

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* remove test_cutedsl_backend.cpp

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* make envvar.rst more human

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

---------

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix python tests

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* skip L1 C++ UT with CUDA when CuTeDSL is not enabled globally

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix C++ UT python discovery

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix deadlock for multigpu

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix python init

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* fix python init again thanks to codex

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* refactor the serialized initial launch

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

* nit

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>

---------

Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Jan Bielak <jbielak@nvidia.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Jan Bielak <janekb04@icloud.com>
Co-authored-by: Jan Bielak <jbielak@nvidia.com>
cael-ling and others added 8 commits September 30, 2026 11:54
…VIDIA#3467)

* [Common] Add single-launch grouped fused amax for row-scaled NVFP4

Compute per-expert rowwise/columnwise amax over a packed (sum_M, K) grouped input in one kernel launch instead of one launch per expert, exposed via nvte_group_nvfp4_compute_amax and gated by NVTE_NVFP4_FUSED_AMAX. Adds a gtest that checks the result against a CPU reference and across launches.

Signed-off-by: Cael Ling <caell@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [Common] Enforce K alignment and per-expert amax buffers in grouped NVFP4 amax

The public entry point calls the launcher directly, bypassing the eligibility
check, so misaligned K silently dropped trailing columns and heterogeneous
per-expert amax buffers could dereference null. Add a column/split alignment
check in the launcher, take do_row/do_col as the union over experts with a
per-expert null guard in the kernel, and cover it with HeterogeneousColumnwise.

Signed-off-by: Cael Ling <caell@nvidia.com>

* [Common] nvfp4: drop unused grouped fused-amax gate and env switch

group_fused_amax_supported was never called (the C API calls the launcher directly and alignment checks already live there), and it carried the NVTE_NVFP4_FUSED_AMAX switch. Remove the dead function and its now-unused <cstdlib>.

Signed-off-by: Cael Ling <caell@nvidia.com>

---------

Signed-off-by: Cael Ling <caell@nvidia.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
NVIDIA#3553)

* Isolate FlashAttention 2/3 imports so a broken install cannot break TE

The version gate only inspects distribution metadata, so a flash-attn whose
extension fails to load -- missing CUDA dependency, ABI mismatch -- raised
straight out of backends.py and made `import transformer_engine.pytorch` fail
outright, even for callers that never ask for FlashAttention.

FlashAttention 4 already guards its import this way, so give FA2 and FA3 the same
treatment: on ImportError leave the module aliases at None, mark the backend
unavailable, and warn with the version and the original error. A genuine TE
failure still raises; only the optional dependency is contained.

The regression test injects a metadata-only flash-attn distribution in a
subprocess, because the fault has to be in place before transformer_engine is
imported. It fails for both FA2 and FA3 without this change and passes with it.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>

* Fix optional FlashAttention import regression coverage

Signed-off-by: Przemek Tredak <ptredak@nvidia.com>

---------

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: Przemek Tredak <ptredak@nvidia.com>
Co-authored-by: Przemek Tredak <ptredak@nvidia.com>
…A#3565)

* [common] Reject mismatched bias/out dtype for non-FP8/FP4 GEMM

cuBLASLt assumes bias dtype equals D when BIAS_DATA_TYPE is unset. TE only
sets that attribute on the FP8/FP4 path, so BF16 bias with FP32 out was
reinterpreted and read past the bias allocation. Match cuBLAS support tables
and Megatron-LM#6000: require bias dtype == D on the non-FP8/FP4 path.

Fixes NVIDIA#3562

Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [qa] Wire test_gemm_bias_dtype into L0 pytorch unittest

L0 enumerates explicit pytest paths, so the new bias dtype regression
module was never invoked by community CI. Add it next to the other
top-level pytorch tests.

Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com>

* [qa] Move non-fp8 bias dtype check into test_sanity

Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com>

---------

Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Przemyslaw Tredak <ptredak@nvidia.com>
Signed-off-by: Santosh Bhavani <santosh.bhavani@live.com>
)

* [PyTorch] MoE Sequential block with Dispatch and Combine as basic ops

Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
Signed-off-by: Varun Thumbe <vthumbe@nvidia.com>

* minor bug fix in test

Signed-off-by: Varun Thumbe <vthumbe@nvidia.com>

---------

Signed-off-by: Varun Thumbe <vthumbe@nvidia.com>
Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
…iled kernel (NVIDIA#3454)

* [Common] nvfp4: fuse row/col amax into a single TMA-tiled kernel

The row-scaled path ran two amax kernels; the columnwise one read global
memory column-major (uncoalesced) and dominated runtime. Compute both
directions in one kernel that streams 128x128 chunks through shared memory
via TMA and reduces columns from SMEM.

Gated by fused_amax_supported (BF16, 128-aligned dims); other cases keep the
two-kernel path. Set NVTE_NVFP4_FUSED_AMAX=0 to force fallback.

Signed-off-by: Cael Ling <caell@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* [Common] Preserve NVFP4 fused amax buffers on no-op quantize

The fused wrapper zeroed both amax buffers with an unconditional memset while
the kernel returns early on noop[0]==1, so a skipped (graph-replay) call cleared
the previously published amax. Replace the memset with a noop-aware zero kernel
that returns early on the same flag, matching the standalone kernels' contract.

Signed-off-by: Cael Ling <caell@nvidia.com>

* [Common] nvfp4: hide row-scaled amax fused/unfused behind one entry

Drop the NVTE_NVFP4_FUSED_AMAX switch and make fused-vs-unfused an internal detail; callers use nvfp4::row_scaled::compute_amaxes, unfused stays as fallback.

Signed-off-by: Cael Ling <caell@nvidia.com>

* Assume fused amax kernel computes both row-wise and col-wise amaxes

Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
Signed-off-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Signed-off-by: Cael Ling <caell@nvidia.com>
Signed-off-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com>
Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants