arXiv Computer Vision

SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis

arXiv:2608. 10519v2 Announce Type: replace Abstract: InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids.

arXiv Computer Vision
Sep 18

Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

The paper investigates how attention sparsity behaves in autoregressive image generation, finding a distinct diagonal sparsity pattern due to spatial locality of visual tokens. It introduces a diagonal‑aware sparse attention mechanism that skips KV entries along the diagonal within a recent window, achieving up to 3.1× higher throughput and 1.19× lower latency with less than 2% quality loss compared to dense inference.

By Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park
arXiv AI
Aug 20

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

The paper introduces SparsePR, a training‑free block‑sparse attention method for video transformers that partitions query‑key responses and reconstructs the residual via probe‑fitted affine corrections. By pairing sampled‑query key responses into K/V groups and using centroids to guide shared routing, SparsePR reduces attention‑reconstruction error across diverse video generation and world‑model tasks. Experiments show consistent error reductions, with probe fitting contributing most of the improvement, while maintaining generation quality at 22.0–26.0% executed‑pair density and delivering 1.48×–2.61× speedups.

By Pardis Taghavi, Reza Langari, Gaurav Pandey
arXiv Machine Learning
Jul 2

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

arXiv:2601. 11641v3 Announce Type: replace-cross Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers to practical deployment.

By Yuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang, Kun Yuan
Hugging Face Trending Papers
Aug 6

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. However, existing approaches predominantly rely on external proxy scorers and rigid heuristic rules, inevitably suffering from misalignment with the target MLLM's intrinsic evidence and failing to accommodate the non-uniform spatiotemporal information density.

arXiv AI
Jul 17

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.

By Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin