The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that generalization error diverges when the number of model parameters approaches a critical value from below.
Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$.
arXiv:2606. 29519v1 Announce Type: new Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data.
By Lorenzo Livi
The paper presents a non-equilibrium dynamical mean-field theory (DMFT) that explains how learning reshapes the dynamics of recurrent neural networks, turning initially chaotic activity into stable, task-dependent behavior. It shows that a slow, feedback-driven learning process gradually increases effective feedback strength, driving the network through a bifurcation that marks the transition from chaotic to stable dynamics. By deriving the two-time correlation function, the authors identify a critical feedback strength and a learning-rate-dependent critical time that separate these regimes, and they demonstrate that the theory accurately predicts the network’s output evolution during training, matching numerical simulations.
By Varun Vaidya
arXiv:2608. 06597v1 Announce Type: cross Abstract: A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention.
By Bj\"orn Ladewig, Ibrahim Talha Ersoy, Karoline Wiesner
The paper investigates gradient descent dynamics in the Edge of Stability regime, where a large learning rate causes persistent oscillations linked to improved generalization. It introduces a tractable continuous‑time mean–fluctuation model that couples the window‑averaged trajectory with its fluctuation covariance, derives this model rigorously from a sharp‑valley framework, and analyzes its stationary states and linear stability. The authors also extend the model to wide two‑layer networks, deriving a Wasserstein‑2 gradient flow for weights and fluctuations, proving well‑posedness, a mean‑field limit, and conditional convergence results, with numerical experiments illustrating the predictions and finite‑time limitations.
By Antonin Chodron de Courcel
The paper investigates the limits of implementing neural network field theory on a computer, focusing on function classes that are regular enough for computation. It presents a no‑go theorem showing that finite‑width network ensembles cannot consistently realize either a quantum or effective field theory due to violations of reflection positivity and lack of scale separation. The study distinguishes between finite‑width and infinite‑width interpretations, concluding that only smeared correlators of the infinite‑width limit are computable with controlled error, and identifies two possible ways to evade the theorem—by relaxing finite variance or exact rotation invariance.
By Thomas R. Harvey
arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv:2606. 17120v1 Announce Type: new Abstract: Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature.
By Ibrahim Talha Ersoy, Karoline Wiesner
arXiv:2401. 04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom.
By Ori Shem-Ur, Yaron Oz
arXiv:2606. 30512v1 Announce Type: cross Abstract: Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theory.
By Srinivasa Rao P., Vangmayi P Reddy
arXiv:2511. 02258v3 Announce Type: replace-cross Abstract: This paper studies the high-dimensional scaling limits of online stochastic gradient descent (SGD).
By Parsa Rangriz