arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.
By Etienne Boursier, Claire Boyer
arXiv:2609.37535v1 Announce Type: new
Abstract: In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps w...
By Sangsidhya Kar
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
By Seyed Morteza Emadi
The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.
By Micah Adler, John W. Byers, Mark Crovella
arXiv:2511.00763v3 Announce Type: replace
Abstract: We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rat...
By Wanda Hou, Leon Zhou, Hong-Ye Hu, Yubei Chen, Yi-Zhuang You, Xiao-Liang Qi
arXiv:2609.37261v1 Announce Type: new
Abstract: Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attent...
By Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone