Infrared Universality of Collective Dynamics across Transformer and State-Space Architectures
arXiv:2608. 18592v1 Announce Type: new Abstract: Whether distinct neural architectures develop common collective dynamics remains an open question.
arXiv:2607. 10923v1 Announce Type: new Abstract: Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood.
arXiv:2608. 18592v1 Announce Type: new Abstract: Whether distinct neural architectures develop common collective dynamics remains an open question.
arXiv:2606. 24396v1 Announce Type: new Abstract: Large Transformer models function as Dense Associative Memories (DAMs), retrieving knowledge via high-dimensional attractor dynamics driven by the self-attention mechanism \citep{ramsauer2020hopfield, wu2024attention}.
arXiv:2610.00423v1 Announce Type: cross Abstract: Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transfor...
The paper investigates associative memory in a bipartite Hopfield–Krotov architecture, termed class H, where hidden neurons serve as the retrieval order parameter. Using the replica method, it derives replica‑symmetric phase diagrams and closed‑form capacities for polynomial load, showing that crosstalk statistics are similar for Ising and spherical visible neurons. With a softmax hidden layer, the load becomes exponential, mapping the thermodynamics onto a random‑energy‑model that exhibits paramagnetic, condensed, and frozen phases, and revealing that heating destabilizes retrieval through quantized attention reassignments while Gaussian patterns remain metastable at all loads.
arXiv:2512.12767v2 Announce Type: replace-cross Abstract: Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation...
arXiv:2608. 12398v1 Announce Type: cross Abstract: We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language.
The paper investigates looped transformers, which perform inference by repeatedly applying a weight‑tied map, making their computation a dynamical process. It shows that identical architectures trained to the same accuracy can converge to distinct dynamical phases—one governed by a saddle‑node fold and another by a Neimark‑Sacker transition—each with different compute scaling behaviors. The study derives a parameter‑free relation linking relaxation time and spectral gap in the fold phase and demonstrates how critical slowing down leads to a workload‑level tail distribution, while the Neimark‑Sacker phase eliminates this scaling law.
arXiv:2607. 15449v1 Announce Type: new Abstract: Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator.
The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.
Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.
arXiv:2609.38768v1 Announce Type: cross Abstract: Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and c...
arXiv:2609.16752v1 Announce Type: new Abstract: Cognitive Field Theory (CFT) proposes that cognition arises from memory-dressed collective dynamics that generate a persistent macroscopic cognitive fi...