arXiv Machine Learning

Attention Function as an Intrinsic Inductive Bias: How Models' Behavior Diverges in Novel Contexts

arXiv Machine Learning
Sep 21

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

The paper investigates how gating the value pathway in attention mechanisms provides two missing capabilities of softmax attention: abstention and noise filtering. Experiments on models ranging from 10M to 350M parameters show that abstention benefits smaller models while noise filtering becomes more advantageous as models scale, and that combining both primitives yields the best performance across all sizes. The authors also demonstrate that the gates effectively suppress interference and that each gate type has a distinct blind spot, all while adding negligible parameters and preserving compatibility with key‑value caching.

By Richard Zhe Wang
arXiv AI
4d ago

Which Attention Heads are like the Human Head? Not the Ones that Compute

The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.

By Christopher Pinier, Gustaw Opie{\l}ka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson
arXiv Machine Learning
Jun 15

Beyond a Single Explanation of the Adam--SGD Gap

arXiv:2606. 14259v1 Announce Type: new Abstract: Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.

By Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni, Jun Pang, Aurelien Lucchi, Antonio Orvieto
arXiv Machine Learning
Sep 21

Gradient-Stable Attention Heads Signal LLM Correctness

The paper introduces HeadEntropy, a training‑free method that predicts the correctness of large language model (LLM) answers by measuring how stable each attention head’s pattern is to further gradient updates. By linking the trace of the softmax Jacobian to 2‑Renyi entropy, the authors show that attention spread correlates with gradient stability, enabling accurate hallucination detection without reference annotations. Across five instruction‑tuned LLMs and five diverse datasets—including medicine, multi‑hop reasoning, and mathematics—HeadEntropy achieves a 0.736 AUROC, outperforming other training‑free baselines and matching hidden‑state probes while incurring less than 1% of inference cost.

By Sophie Ostmeier, Brian Axelrod, Maya Varma, Asad Aali, Yabin Zhang, Magdalini Paschali, Sanmi Koyejo, Curtis Langlotz, Akshay Chaudhari
arXiv Machine Learning
Sep 16

Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation

The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.

By Micah Adler, John W. Byers, Mark Crovella