arXiv Machine Learning

Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality

arXiv Machine Learning
Jul 28

Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere

arXiv:2607. 24502v1 Announce Type: cross Abstract: Rotary position embeddings (RoPE) modify attention scores through position-dependent rotations, but their effect on normalized token dynamics is not captured by the vanilla spherical self-attention model.

By Hao Ye (Xi'an Institute of Optics,Precision Mechanics, Chinese Academy of Sciences, University of Chinese Academy of Sciences)
arXiv AI
Sep 17

The Attention Within: Consensus Dynamics in Selective State Space Models

Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.

By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
arXiv Machine Learning
Sep 4

High-Dimensional Learning Dynamics of Attention-Indexed Models

The paper investigates the training dynamics of attention mechanisms in high-dimensional settings, focusing on attention-indexed models that encompass multi-layer and multi-head architectures. It shows that while the loss landscape can be described by a finite set of trace order parameters, the online stochastic gradient descent dynamics involve an infinite hierarchy of matrix moments that can be accurately approximated by a finite truncated system. The study further reveals that the choice of attention parameterization acts as an implicit bias: untied attention can get trapped in uninformative states, whereas tied attention induces symmetry breaking and enables weak recovery with θ(d² log d) samples, and untied attention exhibits a fast-slow dynamic leading to weak recovery when symmetry is broken.

By Yizhou Xu, Margarita Sagitova, Lenka Zdeborov\'a, Florent Krzakala
arXiv Machine Learning
Jun 9

Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

arXiv:2606. 08985v1 Announce Type: new Abstract: While neural collapse (NC) predicts that a $K$-class-balanced classifier should organize terminal representations as a $(K-1)$-dimensional simplex equiangular tight frame (ETF), modular addition consistently enters a different regime: networks compress to a two-dimensional cyclic geometry in which both classifier weights and token embeddings lie on circles.

By Hu Tan, Kuo Gai, Shihua Zhang
arXiv Machine Learning
Jul 31

Critical attention scaling in long-context transformers

arXiv:2510. 05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length $n$ increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse.

By Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet
arXiv Machine Learning
2d ago

Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

Autoregressive transformers trained on limited trajectories of nonlinear dynamical systems can extrapolate to unseen parameter regimes, reproducing period-doubling cascades, chaotic dynamics, and attractor structures with high fidelity. In the logistic map, the model captures successive period doublings up to period 128, achieving a scaling ratio within $5 imes10^{-4}$ of the Feigenbaum constant. The study also shows how control‑parameter information is processed via attention, shaping the closed‑loop dynamics during training.

By Yilun Liu, Yi Zhang, Ganyu Wu, Sikuan Yan, Mengyue Wang, Alois Knoll, Volker Tresp, Yunpu Ma