arXiv Machine Learning

Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

arXiv:2608. 08922v1 Announce Type: cross Abstract: Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations.

arXiv AI
Sep 17

The Attention Within: Consensus Dynamics in Selective State Space Models

Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.

By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
arXiv Machine Learning
Jul 31

Critical attention scaling in long-context transformers

arXiv:2510. 05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length $n$ increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse.

By Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet
arXiv Machine Learning
Sep 4

High-Dimensional Learning Dynamics of Attention-Indexed Models

The paper investigates the training dynamics of attention mechanisms in high-dimensional settings, focusing on attention-indexed models that encompass multi-layer and multi-head architectures. It shows that while the loss landscape can be described by a finite set of trace order parameters, the online stochastic gradient descent dynamics involve an infinite hierarchy of matrix moments that can be accurately approximated by a finite truncated system. The study further reveals that the choice of attention parameterization acts as an implicit bias: untied attention can get trapped in uninformative states, whereas tied attention induces symmetry breaking and enables weak recovery with θ(d² log d) samples, and untied attention exhibits a fast-slow dynamic leading to weak recovery when symmetry is broken.

By Yizhou Xu, Margarita Sagitova, Lenka Zdeborov\'a, Florent Krzakala