arXiv Machine Learning

Specialization of softmax attention heads: insights from the high-dimensional single-location model

arXiv:2603. 03993v2 Announce Type: replace Abstract: Multi-head attention enables transformer models to represent multiple attention patterns simultaneously.

arXiv Computation and Language
Aug 31

Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs

The paper introduces Semantic Head Specialization (SHS), a phenomenon where Vision Transformer (ViT) attention heads specialize as either object- or background-focused, most evident under full attention. It proposes the SHS-Index to quantify this specialization, demonstrating its ability to distinguish full-attention from chunk-window ViTs and its strong correlation with downstream benchmark performance. Leveraging insights into window interaction, token serialization, and local softmax allocation, the authors design Ariadne Attention, a hybrid attention mechanism that matches full-attention performance on 22 image and video tasks while reducing attention compute by 6.5×.

By Chenhong He, Lei Li, Shicheng Li, Hanglong Lv, Lingpeng Kong, Qi Liu, Tong Yang, Shuhuai Ren
arXiv Machine Learning
Sep 10

Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA

The paper establishes rigorous trainability results for multi-headed attention layers and Low Rank Adaptation (LoRA) models under stochastic training methods. By proving that the empirical regression loss induces a Poincaré inequality with constants independent of data dimension for LoRA and independent of head dimensions for multi-head attention, the authors show that a stochastic differential equation mimicking SGD converges to the loss minima. These results hold without assumptions on data or model size, providing the first theoretical guarantees for training such architectures.

By Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun