arXiv Machine Learning

COBS: Cumulant Order Block Sparse Attention

arXiv:2607. 09052v1 Announce Type: new Abstract: Block sparse attention is a hardware friendly way to alleviate the key-value (KV) cache read bottleneck in large language models (LLMs).

arXiv Machine Learning
4d ago

Block Sparse Flash Attention

Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.

By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv Machine Learning
Sep 23

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV introduces a compensation‑aware sparse attention framework for long‑context LLM inference. It partitions tokens into blocks and optimizes token selection to minimize the error introduced by block‑level mean compensation, using compact block‑level statistics. Experiments on RULER and LongBench‑Pro show CompKV outperforms other sparse baselines and achieves up to a 6.85× speedup over full attention.

By Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu
arXiv Machine Learning
Sep 11

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention proposes a two-stage hierarchical indexer that replaces the flat token scan used in token-level sparse attention mechanisms like DeepSeek Sparse Attention. The method first performs block-level coarse filtering to discard irrelevant regions, then applies the original token-level indexer only within the retained candidate blocks, preserving the same top-sparse pattern for downstream attention. Benchmarks show HISA achieves significant speedups at 64K context and matches the quality of DeepSeek-V3.2 and GLM-5 without additional training.

By Yufei Xu, Fanxu Meng, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Zhaohui Wang, Jiexi Wu, Zhixin Pan, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di Yin, Xing Sun, Muhan Zhang
arXiv AI
Sep 21

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

RBS-Attention introduces a training‑free, radius‑bounded sparse prefill strategy for long‑context large language models, addressing the mean dilution problem where a block centroid can miss highly relevant tokens. The method employs two complementary selection branches: a centroid base branch that captures average relevance and a rescue branch that uses the maximum key‑block radius to flag under‑estimated blocks. Experiments on Qwen3 models demonstrate significant speedups—over 20× in standalone prefill‑attention and nearly 6× in end‑to‑first‑token time—while maintaining competitive accuracy compared to dense attention.

By Chuxu Song, Jiuqi Wei, Zhencan Peng
arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao