arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.
By Etienne Boursier, Claire Boyer
arXiv:2609.37535v1 Announce Type: new
Abstract: In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps w...
By Sangsidhya Kar
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
By Seyed Morteza Emadi
The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.
By Micah Adler, John W. Byers, Mark Crovella
arXiv:2511.00763v3 Announce Type: replace
Abstract: We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rat...
By Wanda Hou, Leon Zhou, Hong-Ye Hu, Yubei Chen, Yi-Zhuang You, Xiao-Liang Qi
arXiv:2609.37261v1 Announce Type: new
Abstract: Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attent...
By Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone
arXiv:2607. 20524v1 Announce Type: new Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval remains untested.
By Sagar Dangal, Manoj Shakya
arXiv:2605. 09778v2 Announce Type: replace Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token.
By Jo\~ao Monteiro, Michal Klein, Pierre Ablin, Marco Cuturi
arXiv:2512.00763v2 Announce Type: replace-cross
Abstract: Adaptive and non-Euclidean optimizers often outperform Euclidean methods such as stochastic gradient descent (SGD) in language modeling by a...
By Robin Yadav, Shuo Xie, Tianhao Wang, Zhiyuan Li
arXiv:2609.25802v1 Announce Type: new
Abstract: We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends onl...
By Moritz Br\"osamle
arXiv:2510. 05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length $n$ increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse.
By Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
By Alexander Hsu, Rongjie Lai