On the Diverse Dynamical Behaviors Arising in Deep Linear Transformers
arXiv:2607. 18584v1 Announce Type: new Abstract: We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems.
We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems. In this perspective, tokens are modeled as particles that interact dynamically through successive linear self-attention layers.
arXiv:2607. 18584v1 Announce Type: new Abstract: We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems.
arXiv:2501. 18322v2 Announce Type: replace Abstract: Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens.
arXiv:2609.37921v1 Announce Type: new Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...
Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.
arXiv:2609.36230v1 Announce Type: new Abstract: We study the dynamical behavior of tokens in transformers from a control-theoretic perspective. Our model includes the feed-forward layer present after...
arXiv:2606. 15207v1 Announce Type: cross Abstract: Transformer architectures have dramatically advanced representation learning and inference in deep models through self-attention mechanisms.
arXiv:2607. 10677v1 Announce Type: new Abstract: Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood.
arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.
Autoregressive transformers trained on limited trajectories of nonlinear dynamical systems can extrapolate to unseen parameter regimes, reproducing period-doubling cascades, chaotic dynamics, and attractor structures with high fidelity. In the logistic map, the model captures successive period doublings up to period 128, achieving a scaling ratio within $5 imes10^{-4}$ of the Feigenbaum constant. The study also shows how control‑parameter information is processed via attention, shaping the closed‑loop dynamics during training.
arXiv:2608. 08922v1 Announce Type: cross Abstract: Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations.
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
arXiv:2609.28448v1 Announce Type: cross Abstract: We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similar...