Conversation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
MultiLatentAttention: low-rank q/kv latents (RMSNorm fused into LayerNormLinear up-projections), decoupled RoPE/NoPE head split with a shared key rope head, DotProductAttention with kv_channels=(qk, v) for the cuDNN fused backend. DeepSeekV3MoE: fused sigmoid router with aux-loss-free expert bias and grouped top-k, routed experts as te.ops GroupedLinear+ScaledSwiGLU+ GroupedLinear (CuTe fused grouped MLP on supported HW), probs applied per-token in the activation, local permute/unpermute or NCCL expert parallelism via ep_dispatch/ep_combine, optional shared expert. DeepSeekV3Layer: pre-RMSNorm + MLA and dense LayerNormMLP (RMSNorm, swiglu) or MoE with residual connections. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
run_deepseek_ep.py checks the EP path against the all-experts-local path numerically (forward, input/gate grads, all-reduced expert wgrads) and smoke-tests the full layer with EP. Also size the default EP recv capacity for per-expert alignment padding and the fused grouped MLP's row-count requirement. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
The per-expert wgrad check called all_reduce on different tensors per rank (rank-local experts), corrupting the reference grads; reduce every expert's grad on every rank instead. Also pass zero-filled recv/grad buffers to ep_dispatch/ep_combine so alignment-padding rows inside the grouped-GEMM m_splits can never poison expert wgrads. Verified on lyris (4x GB300, arm64): run_test_deepseek_ep.sh passes on all ranks (EP forward/dgrad/gate-grad/expert-wgrad match the all-local reference; full-layer EP smoke passes). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Move the Triton MLA RoPE kernels (Megatron-LM fused_mla_yarn_rope_apply port) from tests/pytorch/attention/ mla_rope_utils.py into models/deepseek_v3/mla_rope.py and use them in MultiLatentAttention: the q kernel rotates the rope slice in place and the kv kernel assembles key/value in a single pass, removing the torch.cat/expand/contiguous copies (~10% of layer GPU time). PyTorch fallback (same convention) covers missing Triton and bshd. Fix a latent bug from the test util: the q backward kernel assumed a contiguous incoming gradient, but cuDNN attention backward can hand over a strided one (allocator-state dependent IMA). The old test file stays as a compat shim. Add a Triton-vs-PyTorch parity test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Maps HF DeepseekV3DecoderLayer weights into DeepSeekV3Layer (GLU interleave for routed experts, fused latent norms) and checks forward and input grads match within bf16 tolerance. Expose layernorm_epsilon on MultiLatentAttention (HF latent RMSNorms use 1e-6). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
docs/api/pytorch_models.rst: usage (local and EP), fused-path notes, HF checkpoint weight mapping, and the class API; linked from the PyTorch API page via a toctree entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Keep the verified weight-mapping table in the docs; the comparison itself stays as an out-of-tree script. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…scripts Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…eek_v3.mla_rope directly Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…rical comparison Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…hell launcher Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
… grouped GEMM in local MoE path Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…; expose MLA softmax_scale; clean docstrings Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…ters Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
… routed experts Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…ad of syncing bincount Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…p standard layers, autocast and other utilities Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
DeepSeekV3MoE.make_ep_buffer() builds a buffer for the module's routing config; forward(ep_buffer=...) on the MoE and the layer reuses it instead of creating one per call. The example reuses one buffer by default (--ep-buffer-per-call restores per-call buffers). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
…kdown from an unskewed node All timings are now medians of three runs (spread within 0.3 ms). The TE kernel breakdown is taken from the node not slowed down by nsys, so the rank-skew wait reflects load imbalance rather than profiler skew. Also folds in the pending README restructuring (throughput table, historical naive_grouped profile, results interpretation). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
|
…dates Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
|
Hi @vthumbe1503 @phu0ngng this is PR with DeepSeek layer which uses fused MLP and EP, can you have a look? |
There was a problem hiding this comment.
I wonder if this file should sit together with our other RoPE implementations, i.e. in transformer_engine/pytorch/attention? I feel this models/ directory should contain only the model-level/higher-level implementations.
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Move reusable MLA RoPE kernels and YaRN helpers alongside attention primitives. Update model and MXFP8 attention imports, and move standalone RoPE tests into the attention suite while retaining L0 coverage. Validation: 46 passed, 3 skipped on RTX 5880 Ada using the installed TE primitives; Black, targeted pylint and git diff --check passed. Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Make MLA query RoPE allocate by default and preserve incoming gradients. Keep explicit in-place execution limited to eager autograd. Align MXFP8 and NVFP4 experts to 256 rows independently of fusion, remove the EP tail margin, and cover fused and unfused padding in numerical tests. Simplify model APIs, documentation and the EP benchmark following review. Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
| device = torch.device("cuda", torch.cuda.current_device()) | ||
| return EpBuffer(**self._ep_buffer_kwargs, device=device) | ||
|
|
||
| def _forward_ep(self, tokens: torch.Tensor, ep_buffer=None) -> torch.Tensor: |
There was a problem hiding this comment.
I think it might make sense to use the te.Sequential to define the entire MOE block after this PR gets merged.
#3503
Would be more cleaner from API usage standpoint
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
|
|
||
| GB200 and GB300 GPUs, 4096 tokens per rank, top-k 8, 8 local experts per GPU. | ||
| Times cover one layer's forward + backward; throughput is global, in millions of tokens/s. | ||
| Every number is the median of three independent runs (spread within 0.3 ms). |
There was a problem hiding this comment.
The benchmark results no longer identify when they were measured or which commit was used. That makes the reported timings harder to reproduce or compare with later changes. Please keep the measurement date and commit alongside the results.
| Every number is the median of three independent runs (spread within 0.3 ms). | |
| Every number is the median of three independent runs (spread within 0.3 ms). | |
| Measured on September 30, 2026, at commit | |
| [`939c9db3`](https://github.com/NVIDIA/TransformerEngine/commit/939c9db36afcdb2617392bab57b4249c7cd4bcb0), | |
| before the subsequent merge of `main`. |
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
Signed-off-by: Pawel Gadzinski <pgadzinski@nvidia.com>
| drop_on_overflow = get_ep_drop_on_overflow() | ||
| if drop_on_overflow is None: | ||
| raise RuntimeError("EP requires ep_bootstrap before constructing DeepSeekV3MoE.") |
There was a problem hiding this comment.
Model construction requires bootstrap
If a caller builds a DeepSeekV3MoE with ep_group before calling ep_bootstrap(), this new check raises during construction. Previously, bootstrap was required when creating the EP buffer for forward, so callers could build the model first and bootstrap before execution. That workflow can no longer instantiate the layer.
Description
Adds
transformer_engine.pytorch.models, a namespace for model-specific layers built from TE modules, with a DeepSeek-V3 transformer layer analogous toTransformerLayer.MultiLatentAttention: low-rank q/kv latents with RMSNorm fused into the up-projections, decoupled RoPE/NoPE heads,DotProductAttentionwith asymmetric head dims (cuDNN fused attention), fused Triton MLA RoPE kernels, optional YaRN scaling, TP.DeepSeekV3MoE: fused sigmoid router with aux-loss-free bias and grouped top-k, experts aste.ops.Sequential(GroupedLinear, ScaledSwiGLU, GroupedLinear)(auto-fuses into the CuTe grouped MLP where available), local routing or expert parallelism over NCCL EP, optional shared expert.DeepSeekV3Layer: pre-RMSNorm + MLA, then denseLayerNormMLPorDeepSeekV3MoE, residual connections.Type of change
Changes
transformer_engine/pytorch/models/deepseek_v3/:MultiLatentAttention,DeepSeekV3MoE,DeepSeekV3Layer, MLA RoPE kernels; exported viatransformer_engine.pytorch.modelsModel-specific layerssection indocs/api/pytorch.rsttests/pytorch/test_models.py(RoPE, YaRN, MoE vs dense reference),test_sanity.py::test_sanity_deepseek_v3_layer(all recipes),tests/pytorch/distributed/test_models.py(EP vs all-experts-local, L1)examples/pytorch/deepseek_v3/: runnable EP example with a plain-PyTorch MoE baseline and a README with measurementsPerformance
DeepSeekV3Layerwith DeepSeek-V3 dims, 8 experts per rank, top-k 8, 4096 tokens per rank, fwd+bwd, bf16 unless noted. GB300, 4 GPUs per node, one NVLink domain.naive= torchall_to_all_single+ Python loop of per-expert MLPs,naive_grouped= same all_to_all with a TE grouped GEMM,te=DeepSeekV3MoE. Attention and norms identical in all three.naivenaive_groupedtete, MXFP8te, MXFP8 fused grouped MLPWhere the time goes in the optimized variant (
te, MXFP8 fused, 8 GPUs), per GPU and iteration, fromnsys stats --report cuda_gpu_kern_sum(kernel time 10.8 ms, iteration 10.5 ms without the profiler; no memsets or copies left):nccl_ep_jit_ht_dispatch_kernel(1.05),nccl_ep_jit_ht_combine_kernel(1.29), each twice per iteration (fwd + bwd)local_permute_dup/reduce: staging buffer to expert-major layout, zero-filled paddingncclDevKernel_AllGather_RING_LL(routing-map all-gather in prepare, ~0.06 ms of transfer); the first collective of the layer absorbs load imbalance between ranksgroup_quantize_mxfp8on the recv buffer (0.4),quantize_mxfp8_kernel_cast_onlyfor dense GEMM inputs (0.7)nvjet_sm103_qqtst_*rmsnorm_fwd/bwd(0.39),rotary_*_kv(0.18), residual adds (0.2)For comparison the
naivevariant spends 3.4 ms in all_to_all, 6.5 ms in 8 separate GEMM pairs, and ~12 ms in indexing, zero-fills, copies and adds;naive_groupedremoves the loop (about 10.5 ms) and NCCL EP the remaining all_to_all plus sort / gather / scatter overhead (2.9 ms on 8 GPUs). Full breakdowns inexamples/pytorch/deepseek_v3/README.md.Checklist:
🤖 Generated with Claude Code