arXiv Machine Learning By Sixu Li, Thomas Jacob Maranzatto, Jan Peszek, Trevor Teolis, Semih Akkoc, Konstantin Riedl, Sennur Ulukus, Nicol\'as Garc\'ia Trillos

On the Diverse Dynamical Behaviors Arising in Deep Linear Transformers

Read the original on arXiv Machine Learning →

arXiv:2607. 18584v1 Announce Type: new Abstract: We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
4d ago

Pattern Formation in Transformers

arXiv:2609.37921v1 Announce Type: new Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...

By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
arXiv AI
Sep 17

The Attention Within: Consensus Dynamics in Selective State Space Models

Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.

By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
arXiv Machine Learning
2d ago

Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

Autoregressive transformers trained on limited trajectories of nonlinear dynamical systems can extrapolate to unseen parameter regimes, reproducing period-doubling cascades, chaotic dynamics, and attractor structures with high fidelity. In the logistic map, the model captures successive period doublings up to period 128, achieving a scaling ratio within $5 imes10^{-4}$ of the Feigenbaum constant. The study also shows how control‑parameter information is processed via attention, shaping the closed‑loop dynamics during training.

By Yilun Liu, Yi Zhang, Ganyu Wu, Sikuan Yan, Mengyue Wang, Alois Knoll, Volker Tresp, Yunpu Ma