arXiv:2608. 08922v1 Announce Type: cross Abstract: Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations.
By Qucheng Gao, Zuyi Yang, Xiao Chen
arXiv:2607. 24502v1 Announce Type: cross Abstract: Rotary position embeddings (RoPE) modify attention scores through position-dependent rotations, but their effect on normalized token dynamics is not captured by the vanilla spherical self-attention model.
By Hao Ye (Xi'an Institute of Optics,Precision Mechanics, Chinese Academy of Sciences, University of Chinese Academy of Sciences)
Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.
By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
The paper investigates the training dynamics of attention mechanisms in high-dimensional settings, focusing on attention-indexed models that encompass multi-layer and multi-head architectures. It shows that while the loss landscape can be described by a finite set of trace order parameters, the online stochastic gradient descent dynamics involve an infinite hierarchy of matrix moments that can be accurately approximated by a finite truncated system. The study further reveals that the choice of attention parameterization acts as an implicit bias: untied attention can get trapped in uninformative states, whereas tied attention induces symmetry breaking and enables weak recovery with θ(d² log d) samples, and untied attention exhibits a fast-slow dynamic leading to weak recovery when symmetry is broken.
By Yizhou Xu, Margarita Sagitova, Lenka Zdeborov\'a, Florent Krzakala
arXiv:2609.24202v1 Announce Type: new
Abstract: Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions...
By Jingkun Liu, Yue Song
arXiv:2606. 11585v1 Announce Type: new Abstract: We introduce Kuramoto attention, a self-attention layer in which each hidden coordinate is an angle.
By Joshua Nunley
arXiv:2609. 18145v1 Announce Type: new Abstract: Attention pays, at every layer and for every input, the cost of searching for whom to connect.
By Yoshiaki Takashita
arXiv:2606. 08985v1 Announce Type: new Abstract: While neural collapse (NC) predicts that a $K$-class-balanced classifier should organize terminal representations as a $(K-1)$-dimensional simplex equiangular tight frame (ETF), modular addition consistently enters a different regime: networks compress to a two-dimensional cyclic geometry in which both classifier weights and token embeddings lie on circles.
By Hu Tan, Kuo Gai, Shihua Zhang
arXiv:2608. 18592v1 Announce Type: new Abstract: Whether distinct neural architectures develop common collective dynamics remains an open question.
By Byung Gyu Chae
arXiv:2510. 05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length $n$ increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse.
By Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet
arXiv:2606. 24396v1 Announce Type: new Abstract: Large Transformer models function as Dense Associative Memories (DAMs), retrieving knowledge via high-dimensional attractor dynamics driven by the self-attention mechanism \citep{ramsauer2020hopfield, wu2024attention}.
By Kanishk Awadhiya
Autoregressive transformers trained on limited trajectories of nonlinear dynamical systems can extrapolate to unseen parameter regimes, reproducing period-doubling cascades, chaotic dynamics, and attractor structures with high fidelity. In the logistic map, the model captures successive period doublings up to period 128, achieving a scaling ratio within $5 imes10^{-4}$ of the Feigenbaum constant. The study also shows how control‑parameter information is processed via attention, shaping the closed‑loop dynamics during training.
By Yilun Liu, Yi Zhang, Ganyu Wu, Sikuan Yan, Mengyue Wang, Alois Knoll, Volker Tresp, Yunpu Ma