arXiv:2607. 10923v1 Announce Type: new Abstract: Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood.
By Byung Gyu Chae
Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.
By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
arXiv:2609.28448v1 Announce Type: cross
Abstract: We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similar...
By Qucheng Gao, Zuyi Yang, Xiao Chen
arXiv:2607. 15449v1 Announce Type: new Abstract: Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator.
By Parviz Haggi-Mani, Irina Rish
arXiv:2608. 08922v1 Announce Type: cross Abstract: Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations.
By Qucheng Gao, Zuyi Yang, Xiao Chen
arXiv:2606. 24396v1 Announce Type: new Abstract: Large Transformer models function as Dense Associative Memories (DAMs), retrieving knowledge via high-dimensional attractor dynamics driven by the self-attention mechanism \citep{ramsauer2020hopfield, wu2024attention}.
By Kanishk Awadhiya
The paper investigates looped transformers, which perform inference by repeatedly applying a weight‑tied map, making their computation a dynamical process. It shows that identical architectures trained to the same accuracy can converge to distinct dynamical phases—one governed by a saddle‑node fold and another by a Neimark‑Sacker transition—each with different compute scaling behaviors. The study derives a parameter‑free relation linking relaxation time and spectral gap in the fold phase and demonstrates how critical slowing down leads to a workload‑level tail distribution, while the Neimark‑Sacker phase eliminates this scaling law.
By Gunn Kim
How neural representations preserve the structure of input changes connects representation analysis with internal intervention. We study operable representational content through compatible actions of...
The paper applies variational autoregressive networks to study the time‑dependent joint distribution of the simple exclusion process (SEP) across symmetric, asymmetric, and totally asymmetric variants in one, two, and three dimensions. It reproduces known finite‑time results in 1D and 2D, provides new finite‑time dynamics for all three models, and uncovers scaling relations for the active‑inactive phase transition in 3D, showing a dimension‑independent characteristic length scale. The work offers a unified computational framework for probing nonequilibrium transport dynamics in high‑dimensional configuration spaces.
By Zhimao Liu, Jing Liu, Pan Zhang, Ying Tang
The paper investigates associative memory in a bipartite Hopfield–Krotov architecture, termed class H, where hidden neurons serve as the retrieval order parameter. Using the replica method, it derives replica‑symmetric phase diagrams and closed‑form capacities for polynomial load, showing that crosstalk statistics are similar for Ising and spherical visible neurons. With a softmax hidden layer, the load becomes exponential, mapping the thermodynamics onto a random‑energy‑model that exhibits paramagnetic, condensed, and frozen phases, and revealing that heating destabilizes retrieval through quantized attention reassignments while Gaussian patterns remain metastable at all loads.
By Toshihiro Ota, Masato Taki
The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.
By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv:2512.12767v2 Announce Type: replace-cross
Abstract: Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation...
By Thiparat Chotibut, Oleg Evnin, Weerawit Horinouchi