arXiv AI

LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration

arXiv:2608. 12032v1 Announce Type: cross Abstract: Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and duration grow.

arXiv Computer Vision
Sep 22

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

arXiv:2609.23153v1 Announce Type: new Abstract: Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}:...

By Yuxi Liu, Haoyu Li, Zekun Zhang, Tengxu Sun, Yixiang Cai, Jiayong Li, Yifei Xia, Tianle Liu, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kai Zhang, Kun Yuan, Bin Cui
Hugging Face Trending Papers
Aug 13

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention

Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity.

arXiv Machine Learning
4d ago

Block Sparse Flash Attention

Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.

By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
arXiv Computation and Language
Sep 4

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

By Prakhar Khatri