Hugging Face Trending Papers
Aug 13

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention

Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity.

arXiv Machine Learning
Aug 28

ClusterAttention: A training-free speedup of bidirectional attention

ClusterAttention is a training‑free technique that speeds up bidirectional attention by recursively clustering keys and queries into fixed‑size, power‑of‑two blocks, enabling block‑sparse attention to match dense attention latency on GPUs. The method derives error bounds for sparse attention, showing tighter clusters can reduce error when compensated via centroids, and demonstrates significant speedups—up to six‑fold on large tabular data and 1.8× on video generation—while preserving over 99% of dense accuracy.

By Kasper Nordenram, Amelie Dittmann