arXiv:2608.30720v1 Announce Type: new
Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied...
By Kieran Murphy
arXiv:2410. 24050v3 Announce Type: replace Abstract: Large-scale pretraining of transformers has been central to the success of foundation models.
By Ambroise Odonnat, Wassim Bouaziz, Vivien Cabannes
arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.
By Gregory Duth\'e, Nikolaos Evangelou, Wei Liu, Ioannis G. Kevrekidis, Eleni Chatzi
arXiv:2609.37921v1 Announce Type: new
Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...
By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
arXiv:2609.07086v1 Announce Type: new
Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...
By Hemanth Saratchandran, Simon Lucey
arXiv:2609.08981v1 Announce Type: cross
Abstract: A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing...
By Arman Adibi, Alireza Jafari, Mohammad Ghavamzadeh, Hadi Daneshmand
arXiv:2607. 23869v1 Announce Type: cross Abstract: Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution.
By Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran
arXiv:2606. 29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of deep learning.
By Peilin Liu, Ding-Xuan Zhou
arXiv:2606. 16694v1 Announce Type: cross Abstract: Transformers are widely used as a general-purpose substrate for learning complex correlations between a large collection of coupled variables, but their internal mechanisms have remained mysterious.
By Ravin Raj, Gautam Reddy
The paper introduces Cubit, a Transformer‑style architecture that replaces the standard attention mechanism with Kernel Ridge Regression (KRR). By interpreting attention as Nadaraya‑Watson regression, Cubit incorporates the closed‑form KRR solution, combining kernel‑based value aggregation with normalization via the inverse kernel matrix. The authors also propose a Limited‑Range Rescale (LRR) to stabilize training and report that Cubit shows improved long‑sequence modeling, with gains increasing as training sequence length grows.
By Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Liangchen Tan, Mac Schwager, Anderson Schneider, Yuriy Nevmyvaka, Xiaodong Liu
arXiv:2602. 19143v2 Announce Type: replace Abstract: This paper studies simple transformers trained on a high-order Markov chain, where the model must incorporate information from multiple past positions, each with different statistical importance.
By O\u{g}uz Kaan Y\"uksel, Rodrigo Alvarez Lucendo, Nicolas Flammarion
arXiv:2606. 12058v1 Announce Type: cross Abstract: Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training.
By Itay Lavie, Kirsten Fischer, Andrey Lekov, Frederic Van Maele, Zohar Ringel, Moritz Helias