arXiv Machine Learning

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

The paper investigates how gating the value pathway in attention mechanisms provides two missing capabilities of softmax attention: abstention and noise filtering. Experiments on models ranging from 10M to 350M parameters show that abstention benefits smaller models while noise filtering becomes more advantageous as models scale, and that combining both primitives yields the best performance across all sizes. The authors also demonstrate that the gates effectively suppress interference and that each gate type has a distinct blind spot, all while adding negligible parameters and preserving compatibility with key‑value caching.

arXiv AI
6d ago

Attention Sinks and Outliers in Attention Residuals

The paper introduces OASIS, a method designed to stabilize dual‑normalized attention‑residual architectures by employing explicit null routing and token‑to‑depth null coupling. OASIS mitigates attention sinks and activation outliers, improving low‑bit quantization performance across several language‑model backbones. Empirical results show significant reductions in attention norms and perplexity, with notable gains on long‑context benchmarks.

By Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen, Eric Hanchen Jiang, Yijiang Li, Chenghao Qiu, Chenwei Xu, Zhenyu Pan, Haotian Zhang, Binghui Wang, Yan Chen
arXiv Machine Learning
Jun 2

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

arXiv:2606. 01923v1 Announce Type: cross Abstract: Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations.

By Mingkuan Zhao, Yide Gao, Wentao Hu, Suquan Chen, Tianchen Huang, Zhenhua An, Zetao Chang, Xiayu Sun, Yuheng Min
arXiv AI
Sep 10

Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

The paper investigates whether recent attention‑mechanism improvements—specifically gated attention, Kimi K3, Kimi Delta Attention, and Attention Residuals—effectively eliminate the attention‑sink problem when scaling language models to a one‑million‑token context window. Using a new diagnostic suite called SinkProbe, the authors evaluate sink mass, massive activation, position‑resolved recall, and the recency gap across four small models that vary only in token mixing and depth. Their findings show that the training objective, rather than the architecture, drives the emergence of attention sinks; gating did not replicate its previously reported benefits at the larger scale, and sink mass, activations, and positional bias behaved independently.

By Sara Rizwan, Samaanah Abdus Salam
arXiv AI
Jun 9

Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications

arXiv:2603. 22473v2 Announce Type: replace-cross Abstract: Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the functional contribution of each component type remains insufficiently characterized.

By Hector Borobia, Elies Segu\'i-Mas, Guillermina Tormo-Carb\'o