arXiv:2606. 10384v1 Announce Type: cross Abstract: Criticality has been proposed as a key organizing principle in biological neural systems, yet its origin and relevance in artificial neural networks remain unclear.
By Feixiang Ren, Ling Feng
The paper presents a non-equilibrium dynamical mean-field theory (DMFT) that explains how learning reshapes the dynamics of recurrent neural networks, turning initially chaotic activity into stable, task-dependent behavior. It shows that a slow, feedback-driven learning process gradually increases effective feedback strength, driving the network through a bifurcation that marks the transition from chaotic to stable dynamics. By deriving the two-time correlation function, the authors identify a critical feedback strength and a learning-rate-dependent critical time that separate these regimes, and they demonstrate that the theory accurately predicts the network’s output evolution during training, matching numerical simulations.
By Varun Vaidya
arXiv:2607. 21716v1 Announce Type: new Abstract: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics.
By Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian
arXiv:2606. 30930v1 Announce Type: cross Abstract: Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory.
By Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun, Rene Vidal
arXiv:2607. 27731v1 Announce Type: new Abstract: Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics.
By Jiaxiang Li, Zhiqi Bu, Shiyun Xu
arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
By Alessandro Ronca
arXiv:2606. 15551v1 Announce Type: new Abstract: The Edge of Stability (EoS) phenomenon, where gradient descent operates with sharpness exceeding the classical convergence threshold yet the loss decreases over long timescales, is ubiquitous in modern deep learning but remains poorly understood in realistic settings.
By Eric Gan
arXiv:2606. 09929v1 Announce Type: cross Abstract: Physical reservoir computing harnesses nonlinear mechanical dynamics but, by convention, freezes the substrate and trains only a linear readout, presuming the substrate is not usefully trainable.
By Caleb Munigety
The paper studies how Adam’s two momentum timescales, β1 and β3, influence loss spikes during neural‑network training. By mapping training dynamics across the (β1,β3) plane, the authors find an approximately linear boundary, 1-β3 = C(1-β1), that separates spiky from non‑spiky behavior, with the coefficient C linked to the effective loss exponent in superquadratic loss functions. They also show that confident cross‑entropy losses create a core–wall landscape that behaves superquadratically at the scale of an optimizer update, explaining the observed spikes.
By Gaoxiang Tang, Huanran Chen, Ziming Liu
arXiv:2508. 03105v3 Announce Type: replace Abstract: We analyze the convergence behavior of stochastic gradient descent with momentum (SGDM) under dynamic learning-rate and batch-size schedules by introducing a novel and simpler Lyapunov function.
By Yuichi Kondo, Hideaki Iiduka
arXiv:2607. 21005v1 Announce Type: new Abstract: Most explanations of training instability focus on \emph{learning-rate criticality}, typically characterized by the Edge of Stability, beyond which optimization becomes unstable.
By Xiaolong Li, Zhangchen Zhou, Zhi-Qin John Xu
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento