arXiv:2609.01311v1 Announce Type: new
Abstract: We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multicla...
By Skanda Athreya, Yutong Wang
arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.
By Phokion Kolaitis, Rik Sengupta
arXiv:2602.05896v3 Announce Type: replace-cross
Abstract: Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in...
By Alexander Kozachinskiy, Tomasz Steifer, Przemys{\l}aw Wa{\l}\c{e}ga
arXiv:2410. 11500v2 Announce Type: replace-cross Abstract: In this paper, we establish a collection of covering number bounds for linear function classes under various norm constraints on the inputs and matrices.
By Lan V. Truong
arXiv:2605. 22223v2 Announce Type: replace Abstract: We study how we can leverage only a handful of characteristics of a transformer's architecture to closely predict the number of different sequences it can output, both qualitatively and quantitatively.
By Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan
arXiv:2609.36698v1 Announce Type: new
Abstract: To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studi...
By Ene Meco, Emadeldeen Hamdan, A. Enis Cetin