Attention Function as an Intrinsic Inductive Bias: How Models' Behavior Diverges in Novel Contexts
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates how gating the value pathway in attention mechanisms provides two missing capabilities of softmax attention: abstention and noise filtering. Experiments on models ranging from 10M to 350M parameters show that abstention benefits smaller models while noise filtering becomes more advantageous as models scale, and that combining both primitives yields the best performance across all sizes. The authors also demonstrate that the gates effectively suppress interference and that each gate type has a distinct blind spot, all while adding negligible parameters and preserving compatibility with key‑value caching.
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
arXiv:2508. 17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms.
The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms.
arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.