arXiv Machine Learning By Thomas Rossi

Uncertainty-gated selection for block-sparse attention

Read the original on arXiv Machine Learning →

arXiv:2607. 07724v1 Announce Type: new Abstract: Block-sparse attention scales long-context language models by replacing the O(N^2) softmax with a per-query top-k selection over key blocks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
arXiv AI
Sep 21

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

RBS-Attention introduces a training‑free, radius‑bounded sparse prefill strategy for long‑context large language models, addressing the mean dilution problem where a block centroid can miss highly relevant tokens. The method employs two complementary selection branches: a centroid base branch that captures average relevance and a rescue branch that uses the maximum key‑block radius to flag under‑estimated blocks. Experiments on Qwen3 models demonstrate significant speedups—over 20× in standalone prefill‑attention and nearly 6× in end‑to‑first‑token time—while maintaining competitive accuracy compared to dense attention.

By Chuxu Song, Jiuqi Wei, Zhencan Peng
arXiv Machine Learning
4d ago

Block Sparse Flash Attention

Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.

By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata