[PR 1/2] SVG implementation for LTX 2 - #497
Open
jitendra-jalwaniya wants to merge 1 commit into
Open
jitendra-jalwaniya wants to merge 1 commit into
jitendra-jalwaniya wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Code Review
This pull request integrates Sparse VideoGen (SVG) attention into the LTX2 model. It introduces SVG configuration parameters, updates the attention layer to route between dense and sparse SVG attention based on active steps and layers, and propagates the necessary spatiotemporal and step metadata through the transformer blocks and static/block contexts. Additionally, comprehensive unit tests are added to verify the SVG activation boundaries, dispatch routing, and full model forward passes. There are no review comments, so no additional feedback is provided.
jitendra-jalwaniya
force-pushed
the
ltx2_block_benchmark_fixes
branch
from
September 29, 2026 07:41
ecc5b5e to
55e6847
Compare
jitendra-jalwaniya
force-pushed
the
ltx2_svg_model
branch
from
September 29, 2026 07:41
6e1b301 to
ee0e722
Compare
jitendra-jalwaniya
requested review from
Perseus14
and removed request for
entrpn
September 29, 2026 07:55
jitendra-jalwaniya
force-pushed
the
ltx2_svg_model
branch
from
September 29, 2026 10:36
ee0e722 to
650b1c6
Compare
jitendra-jalwaniya
force-pushed
the
ltx2_block_benchmark_fixes
branch
from
September 29, 2026 10:36
55e6847 to
2196275
Compare
jitendra-jalwaniya
force-pushed
the
ltx2_svg_model
branch
from
September 29, 2026 12:36
650b1c6 to
9b30925
Compare
jitendra-jalwaniya
changed the base branch from
ltx2_block_benchmark_fixes
to
fix/pyink-main
September 29, 2026 17:42
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR extends Sparse VideoGen (SVG) spatiotemporal attention support to LTX-2 (LTX2) video generation models on Cloud TPUs, building on the custom Ulysses/ring SVG kernel infrastructure introduced for Wan (PR #480).
Self-attention in LTX-2 transformer blocks dynamically profiles query tokens to choose between spatial and temporal attention patterns per head, skipping unneeded query–key interactions while executing through hardware-aligned local-band kernels on TPU. Sparse attention is opt-in (
use_svg_attention: True), disabled by default, and configurable across denoising steps, layers, and sparsity densities. Audio self-attention and cross-modal attention remain dense to preserve temporal and semantic grounding.This is PR 1/2 (model side). It depends on #493 (pyink formatting fix on
main). The config, pipeline, AOT metadata and docs wiring are in #498.Changes in this PR:
attention_ltx2.py:LTX2Attentionaccepts anattention_configdict with SVG settings and dispatches video self-attention to the SVG kernel (or dense viajax.lax.cond) based on the active step/layer window.transformer_ltx2.py: plumbsspatiotemporal_shape,svg_timestep,svg_step_indexand per-layerlayer_indexthroughLTX2StaticContext/LTX2BlockContext(scanned and unscanned paths). Onlyattn1(video self-attention) gets SVG;audio_attn1is forced dense.tests/ltx2/test_svg_attention_ltx2.py: new unit tests.VABench Evaluation: SVG vs. Dense Attention
The end-to-end results below require both this PR and #498.
We evaluated SVG against dense attention on the Full VABench Benchmark suite (778 prompts across all 24 Easy/Hard bundles and 7 content categories) for LTX-2 synchronized text-to-audio-video (T2AV) generation at long sequence length (768 × 1280 × 241 frames,$N = 29,760$ video tokens, 10.04s @ 24 fps video + 24 kHz PCM audio) on TPU v6e-8 (8 chips), followed by a 15-dimension VABench evaluation across 8× NVIDIA A100-80GB GPUs. Each prompt was generated once per configuration:
use_svg_attention=False,attention=ulysses_customuse_svg_attention=True,svg_spatial_density=0.25,attention=ulysses_custom,svg_active_*left at defaults (SVG active on all steps and layers)1. TPU v6e-8 Generation Performance (
778 Videos @ 768 × 1280 × 241)2. 15-Dimension VABench Quality Highlights
Enabling SVG yields faster generation with comparable overall quality: most metrics are on par or slightly higher, with a small drop in judged visual realism (-1.78%):
second_desyncsecond_lsaQwen2.5-Omni-7B):Full 15-Dimension VABench Comparison Table (
778 Prompts)first_dnsmossig_bak_ovr+p808)first_nisqafirst_audioboxsecond_viclipsecond_clapsecond_imagebindsecond_desyncsecond_lsaQwen2.5-Omni-7B)third_alignmentthird_audio_realitythird_visual_realitythird_expressivenessthird_artistryfourth_qa_audiofourth_qa_visionTesting
Run the LTX-2 SVG attention unit tests from the repository root:
All existing Wan and LTX-2 unit tests continue to pass.