A Unified Perspective on the Dynamics of Deep Transformers
arXiv:2501. 18322v2 Announce Type: replace Abstract: Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens.
arXiv:2606. 07600v1 Announce Type: cross Abstract: We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability measures.
arXiv:2501. 18322v2 Announce Type: replace Abstract: Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens.
arXiv:2607. 27975v1 Announce Type: new Abstract: We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions.
We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem.
arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.
The paper studies universality in non‑separable Approximate Message Passing (AMP) algorithms. It introduces a Bounded Composition Property (BCP) for polynomial non‑linearities and a BCP‑approximability condition for Lipschitz AMP, showing that these conditions guarantee state‑evolution universality for matrices with non‑Gaussian entries. The authors demonstrate that many common non‑separable non‑linearities—such as local denoisers, spectral denoisers, and compositions of separable functions with generic linear maps—satisfy these conditions, thereby extending universality results beyond Gaussian or rotationally‑invariant data.
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
arXiv:2604. 24196v4 Announce Type: replace-cross Abstract: A drifting model is a one-step generator trained by moving each sample along a field of kernel-weighted attraction toward data samples and repulsion between model samples; training halts once this field vanishes.
arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.
arXiv:2608. 02487v1 Announce Type: cross Abstract: Recently, rectified flow has emerged as a fundamental framework for large-scale image generation, powering state-of-the-art systems such as FLUX.
arXiv:2608. 12828v1 Announce Type: cross Abstract: Distribution steering seeks feedback laws that drive the state law of a dynamical system between prescribed initial and terminal distributions.
arXiv:2608. 04531v1 Announce Type: new Abstract: Functional flow matching is posed on distributions of functions but implemented from finitely many coefficients or point values.
arXiv:2607. 18584v1 Announce Type: new Abstract: We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems.