Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity.
arXiv:2607. 03012v1 Announce Type: cross Abstract: Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation.
By Dongyeun Lee, Amir Zandieh, Vahab Mirrokni, Junmo Kim, Insu Han
arXiv:2602. 01801v2 Announce Type: replace-cross Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines.
By Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
arXiv:2609.37001v1 Announce Type: cross
Abstract: Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the compu...
By Xingyu Jia, Baole Ai, Ang Wang, Kang Zhao, Yong Li
arXiv:2609.15810v1 Announce Type: new
Abstract: Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost,...
By Xingyang Li, Dongyun Zou, Shining Zhang, Jiacheng Chen, Haocheng Xi, Lvmin Zhang, Jun-Yan Zhu, Song Han, Zhekai Zhang, Yujun Lin, Muyang Li
ClusterAttention is a training‑free technique that speeds up bidirectional attention by recursively clustering keys and queries into fixed‑size, power‑of‑two blocks, enabling block‑sparse attention to match dense attention latency on GPUs. The method derives error bounds for sparse attention, showing tighter clusters can reduce error when compensated via centroids, and demonstrates significant speedups—up to six‑fold on large tabular data and 1.8× on video generation—while preserving over 99% of dense accuracy.
By Kasper Nordenram, Amelie Dittmann
The paper introduces SparsePR, a training‑free block‑sparse attention method for video transformers that partitions query‑key responses and reconstructs the residual via probe‑fitted affine corrections. By pairing sampled‑query key responses into K/V groups and using centroids to guide shared routing, SparsePR reduces attention‑reconstruction error across diverse video generation and world‑model tasks. Experiments show consistent error reductions, with probe fitting contributing most of the improvement, while maintaining generation quality at 22.0–26.0% executed‑pair density and delivering 1.48×–2.61× speedups.
By Pardis Taghavi, Reza Langari, Gaurav Pandey
arXiv:2607. 25266v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible.
By Ghazal Kaviani, Ghassan AlRegib
arXiv:2607. 27692v1 Announce Type: cross Abstract: Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries.
By Wenshuai Yao, Wenyong Zhou, Hanyong Shao, Yizhe Chen, Zhiyuan Ning, Yuannuo Feng, Ru Huang, Kechao Tang
arXiv:2608. 10519v2 Announce Type: replace Abstract: InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids.
By Jongbeom Lee, Hyunwoo Yu, Jincheol Yang, Jaemin Choi, Suk-Ju Kang
InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context, however, leave late-scale attention costly and make sparse patterns reused from diffusion or image VAR models unreliable.
arXiv:2604. 00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens.
By Dong-Jae Lee, Sunghyun Baek, Junmo Kim