arXiv AI By Pardis Taghavi, Reza Langari, Gaurav Pandey

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

Read the original on arXiv AI →

The paper introduces SparsePR, a training‑free block‑sparse attention method for video transformers that partitions query‑key responses and reconstructs the residual via probe‑fitted affine corrections. By pairing sampled‑query key responses into K/V groups and using centroids to guide shared routing, SparsePR reduces attention‑reconstruction error across diverse video generation and world‑model tasks. Experiments show consistent error reductions, with probe fitting contributing most of the improvement, while maintaining generation quality at 22.0–26.0% executed‑pair density and delivering 1.48×–2.61× speedups.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 13

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention

Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity.

arXiv AI
Jul 17

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.

By Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin
arXiv Machine Learning
Aug 28

ClusterAttention: A training-free speedup of bidirectional attention

ClusterAttention is a training‑free technique that speeds up bidirectional attention by recursively clustering keys and queries into fixed‑size, power‑of‑two blocks, enabling block‑sparse attention to match dense attention latency on GPUs. The method derives error bounds for sparse attention, showing tighter clusters can reduce error when compensated via centroids, and demonstrates significant speedups—up to six‑fold on large tabular data and 1.8× on video generation—while preserving over 99% of dense accuracy.

By Kasper Nordenram, Amelie Dittmann