arXiv:2501. 18322v2 Announce Type: replace Abstract: Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens.
By Val\'erie Castin, Pierre Ablin, Jos\'e Antonio Carrillo, Gabriel Peyr\'e
arXiv:2607. 27975v1 Announce Type: new Abstract: We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions.
By Ka\u{g}an Akman, Naci Saldi, Serdar Y\"uksel
We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem.
arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.
By Gregory Duth\'e, Nikolaos Evangelou, Wei Liu, Ioannis G. Kevrekidis, Eleni Chatzi
The paper studies universality in non‑separable Approximate Message Passing (AMP) algorithms. It introduces a Bounded Composition Property (BCP) for polynomial non‑linearities and a BCP‑approximability condition for Lipschitz AMP, showing that these conditions guarantee state‑evolution universality for matrices with non‑Gaussian entries. The authors demonstrate that many common non‑separable non‑linearities—such as local denoisers, spectral denoisers, and compositions of separable functions with generic linear maps—satisfy these conditions, thereby extending universality results beyond Gaussian or rotationally‑invariant data.
By Max Lovig, Tianhao Wang, Zhou Fan
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
By Alexander Hsu, Rongjie Lai