The paper investigates how gating the value pathway in attention mechanisms provides two missing capabilities of softmax attention: abstention and noise filtering. Experiments on models ranging from 10M to 350M parameters show that abstention benefits smaller models while noise filtering becomes more advantageous as models scale, and that combining both primitives yields the best performance across all sizes. The authors also demonstrate that the gates effectively suppress interference and that each gate type has a distinct blind spot, all while adding negligible parameters and preserving compatibility with key‑value caching.
By Richard Zhe Wang
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
By Subham Singh, Ashutosh Mishra, Subha Raut
arXiv:2508. 17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms.
By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.
By Christopher Pinier, Gustaw Opie{\l}ka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms.
arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.
By Etienne Boursier, Claire Boyer
arXiv:2606. 12058v1 Announce Type: cross Abstract: Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training.
By Itay Lavie, Kirsten Fischer, Andrey Lekov, Frederic Van Maele, Zohar Ringel, Moritz Helias
arXiv:2606. 14259v1 Announce Type: new Abstract: Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.
By Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni, Jun Pang, Aurelien Lucchi, Antonio Orvieto
The paper introduces HeadEntropy, a training‑free method that predicts the correctness of large language model (LLM) answers by measuring how stable each attention head’s pattern is to further gradient updates. By linking the trace of the softmax Jacobian to 2‑Renyi entropy, the authors show that attention spread correlates with gradient stability, enabling accurate hallucination detection without reference annotations. Across five instruction‑tuned LLMs and five diverse datasets—including medicine, multi‑hop reasoning, and mathematics—HeadEntropy achieves a 0.736 AUROC, outperforming other training‑free baselines and matching hidden‑state probes while incurring less than 1% of inference cost.
By Sophie Ostmeier, Brian Axelrod, Maya Varma, Asad Aali, Yabin Zhang, Magdalini Paschali, Sanmi Koyejo, Curtis Langlotz, Akshay Chaudhari
The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.
By Micah Adler, John W. Byers, Mark Crovella
arXiv:2603. 03993v2 Announce Type: replace Abstract: Multi-head attention enables transformer models to represent multiple attention patterns simultaneously.
By M. Sagitova, O. Duranthon, L. Zdeborov\'a
arXiv:2508. 08289v3 Announce Type: replace Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave.
By Mu Qiao