arXiv Machine Learning

What Makes Position Zero Special? A Mechanistic Study of Position Zero Attention Sinks in LLMs

arXiv:2603. 06591v2 Announce Type: replace Abstract: Transformers frequently allocate disproportionate attention to specific tokens, a phenomenon known as attention sinks.

arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv AI
Sep 10

Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

The paper investigates whether recent attention‑mechanism improvements—specifically gated attention, Kimi K3, Kimi Delta Attention, and Attention Residuals—effectively eliminate the attention‑sink problem when scaling language models to a one‑million‑token context window. Using a new diagnostic suite called SinkProbe, the authors evaluate sink mass, massive activation, position‑resolved recall, and the recency gap across four small models that vary only in token mixing and depth. Their findings show that the training objective, rather than the architecture, drives the emergence of attention sinks; gating did not replicate its previously reported benefits at the larger scale, and sink mass, activations, and positional bias behaved independently.

By Sara Rizwan, Samaanah Abdus Salam
arXiv Machine Learning
Aug 31

Sliding-window beats linear attention

arXiv:2608.28444v1 Announce Type: cross Abstract: Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previo...

By Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais
arXiv Machine Learning
Sep 23

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV introduces a compensation‑aware sparse attention framework for long‑context LLM inference. It partitions tokens into blocks and optimizes token selection to minimize the error introduced by block‑level mean compensation, using compact block‑level statistics. Experiments on RULER and LongBench‑Pro show CompKV outperforms other sparse baselines and achieves up to a 6.85× speedup over full attention.

By Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu
arXiv Machine Learning
Sep 18

Post-Boundary Bridge: Must Local Attention Go Global Between Global Layers?

The paper introduces Post-Boundary Bridge (PBB), a hybrid transformer architecture that keeps causal attention within blocks while adding direct connections across block boundaries. PBB focuses on within-block modeling and nearby exchange, delegating long-range communication to full-attention layers. Experiments on dense and mixture-of-experts models ranging from 205 million to 2.07 billion parameters show that PBB hybrids maintain near-full perplexity, competitive downstream performance, and improved source retrieval, while Flash-PBB achieves 1.82× faster decoding with half the local key‑value cache compared to Flash‑SWA.

By Zhibo Yang
arXiv AI
Jul 22

A Controlled Study of Attention-Only Transformers

arXiv:2607. 18363v1 Announce Type: cross Abstract: Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once.

By Henry Ndubuaku, Karen Mosoyan, Jakub Mroz, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H Lee
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen
arXiv Machine Learning
Jul 10

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

arXiv:2601. 12145v3 Announce Type: replace Abstract: Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase.

By Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao