arXiv AI By Marios Papamichalis, Regina Ruane

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

Read the original on arXiv AI →

arXiv:2608. 14712v1 Announce Type: cross Abstract: Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Which Attention Heads are like the Human Head? Not the Ones that Compute

The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.

By Christopher Pinier, Gustaw Opie{\l}ka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson
arXiv Machine Learning
Sep 25

Invertible Query-Key Coupling Composes with Attention Mechanisms

The paper introduces a coupled query‑key transformation that jointly evolves queries and keys via an invertible coupling before the standard dot‑product scoring in attention mechanisms. Implemented as a lightweight alternating affine map, the coupling is added on top of existing attention methods and preserves the original softmax and architecture. Experiments on WikiText‑103 show that coupling improves performance when combined with Differential Attention, query‑key normalization, and Multi‑Token Attention, especially at larger model scales, while its standalone benefit diminishes with size.

By Barak Gahtan, Alex M. Bronstein