The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelength limits how far it can discriminate position. Aligned with this structure, we propose the per-RoPE-wavelength distance window: it prunes the query--key inner-product terms beyond a wavelength-proportional distance.
arXiv:2608. 06849v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs.
By Yehan Yang, Junyuan Shang, Yang Li, Guanqun Zhao, Shuohuan Wang, Dianhai Yu
Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.
By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv:2607. 24555v1 Announce Type: cross Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step.
By Junsung Hwang
arXiv:2604. 00004v2 Announce Type: replace-cross Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT).
By Ning Yang, Hengyu Zhong, Wentao Wang, Baoliang Tian, Haijun Zhang, Jun Wang
arXiv:2609.30738v1 Announce Type: new
Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score,...
By Tianfang Xie, Wei Zhu
arXiv:2607. 19368v1 Announce Type: new Abstract: Long-prompt inference remains expensive because prefill attention scales quadratically with sequence length.
By Ali Mahdavi, Azaseh Zamanifar, Amirfarhad Farhadi, Omid Kashefi
arXiv:2607. 07740v1 Announce Type: cross Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints.
By Haozhan Tang, Zerui Wang, Yuxian Gu, Song Han, Han Cai
arXiv:2606. 01294v1 Announce Type: cross Abstract: Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tasks.
By Dong Le, Thong Nguyen, Cong-Duy Nguyen, Anh Tuan Luu
arXiv:2607. 27692v1 Announce Type: cross Abstract: Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries.
By Wenshuai Yao, Wenyong Zhou, Hanyong Shao, Yizhe Chen, Zhiyuan Ning, Yuannuo Feng, Ru Huang, Kechao Tang
arXiv:2607. 07724v1 Announce Type: new Abstract: Block-sparse attention scales long-context language models by replacing the O(N^2) softmax with a per-query top-k selection over key blocks.
By Thomas Rossi
NestedKV is a training‑free key‑only KV cache compression technique for long‑context language models that uses global, block‑level, and sliding‑window key anchors to score tokens via multi‑time‑scale cosine anomaly. It combines these rankings with a head‑adaptive outer learner and surprise‑gated token routing, requiring no model modification or additional training. Experiments on Qwen3 and Llama‑3.2 across benchmarks such as RULER, LongBench, and MMLU‑Pro show that NestedKV outperforms existing methods when the retained cache is small, achieving up to 19‑point gains on RULER and LongBench at a 75% retention rate.
By Hong Chen, Xiang Liu, Yubo Gao, Yuxuan Fan, Bo Wang, Yuanlin Chu, Yuanguo Lin, Xuming Hu