arXiv Machine Learning

Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective

arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.

arXiv Machine Learning
1d ago

Universal interpolation for deep residual self-attention networks

The paper proves that deep residual self‑attention networks can universally interpolate between any two collections of sequences using only two fixed single‑head attention blocks with Gaussian‑initialized projections. The interpolation is achieved by varying the order, signs, and durations of these blocks, independent of the specific input and output sequences. The result holds for both continuous and finite depth, and the authors also extend the analysis to causal‑masked settings.

By Sibylle Marcotte, Joan Bruna
arXiv Machine Learning
Jun 9

Token Sample Complexity of Attention

arXiv:2512. 10656v3 Announce Type: replace Abstract: As context windows in large language models continue to expand, it is essential to characterize how attention behaves at extreme sequence lengths.

By L\'ea Bohbot, Cyril Letrouit, Gabriel Peyr\'e, Fran\c{c}ois-Xavier Vialard
arXiv Machine Learning
Jul 10

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

arXiv:2601. 12145v3 Announce Type: replace Abstract: Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase.

By Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao
arXiv AI
3d ago

Switching Linear Attention

Switching Linear Attention (SwiLA) is a new sequence layer that improves upon standard softmax attention by maintaining a fixed-size recurrent state while enhancing representational capacity. It derives its recurrence from a test-time regression framework, using online expectation-maximization in a mixture of linear regressions model. In various benchmarks—including associative recall, in-context language learning, and language modeling—SwiLA achieves strong performance, narrowing the gap to softmax attention and even surpassing it in some settings.

By Hyun Dong Lee, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, Scott W. Linderman