arXiv:2609.01311v1 Announce Type: new
Abstract: We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multicla...
By Skanda Athreya, Yutong Wang
arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.
By Phokion Kolaitis, Rik Sengupta
arXiv:2602.05896v3 Announce Type: replace-cross
Abstract: Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in...
By Alexander Kozachinskiy, Tomasz Steifer, Przemys{\l}aw Wa{\l}\c{e}ga
arXiv:2410. 11500v2 Announce Type: replace-cross Abstract: In this paper, we establish a collection of covering number bounds for linear function classes under various norm constraints on the inputs and matrices.
By Lan V. Truong
arXiv:2605. 22223v2 Announce Type: replace Abstract: We study how we can leverage only a handful of characteristics of a transformer's architecture to closely predict the number of different sequences it can output, both qualitatively and quantitatively.
By Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan
arXiv:2609.36698v1 Announce Type: new
Abstract: To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studi...
By Ene Meco, Emadeldeen Hamdan, A. Enis Cetin
The paper introduces fixed universal transformers, which are transformers with immutable internal parameters that can emulate any transformer within a specified class by encoding the target model’s description into the input embedding. The authors provide explicit sparse constructions that achieve universality when the embedding dimension is large enough, and demonstrate that universality is generic—randomly initialized transformers are almost surely universal. Empirical tests on parenthesis balancing and multi‑hop reasoning tasks support the theory, suggesting that a transformer’s expressive power largely stems from its input representation rather than its learned weights.
By Jingwen Liu, Alexandr Andoni, Daniel Hsu
arXiv:2606. 08768v1 Announce Type: new Abstract: Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings.
By Blanka K\"over, Alexandra Butoi, Anej Svete, Michael Hahn, Ryan Cotterell
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
By Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis
arXiv:2603. 02238v2 Announce Type: replace Abstract: Length generalization is a key property of a learning algorithm that enables it to make correct predictions on inputs of any length, given finite training data.
By Andy Yang, Pascal Bergstr\"a{\ss}er, Georg Zetzsche, David Chiang, Anthony W. Lin
arXiv:2603. 19954v2 Announce Type: replace Abstract: Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected has been limited.
By Yash Sarrof, Yupei Du, Katharina Stein, Alexander Koller, Sylvie Thi\'ebaux, Michael Hahn