arXiv Machine Learning

Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

arXiv:2607. 10923v1 Announce Type: new Abstract: Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood.

arXiv Machine Learning
Sep 11

Phases in a class of associative memories via hidden neurons

The paper investigates associative memory in a bipartite Hopfield–Krotov architecture, termed class H, where hidden neurons serve as the retrieval order parameter. Using the replica method, it derives replica‑symmetric phase diagrams and closed‑form capacities for polynomial load, showing that crosstalk statistics are similar for Ising and spherical visible neurons. With a softmax hidden layer, the load becomes exponential, mapping the thermodynamics onto a random‑energy‑model that exhibits paramagnetic, condensed, and frozen phases, and revealing that heating destabilizes retrieval through quantized attention reassignments while Gaussian patterns remain metastable at all loads.

By Toshihiro Ota, Masato Taki
arXiv Machine Learning
Aug 28

Dynamical phase selection controls compute scaling in looped transformers

The paper investigates looped transformers, which perform inference by repeatedly applying a weight‑tied map, making their computation a dynamical process. It shows that identical architectures trained to the same accuracy can converge to distinct dynamical phases—one governed by a saddle‑node fold and another by a Neimark‑Sacker transition—each with different compute scaling behaviors. The study derives a parameter‑free relation linking relaxation time and spectral gap in the fold phase and demonstrates how critical slowing down leads to a workload‑level tail distribution, while the Neimark‑Sacker phase eliminates this scaling law.

By Gunn Kim
arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv AI
Sep 17

The Attention Within: Consensus Dynamics in Selective State Space Models

Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.

By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada