arXiv AI

Mean--Fluctuation Dynamics at the Edge of Stability

The paper investigates gradient descent dynamics in the Edge of Stability regime, where a large learning rate causes persistent oscillations linked to improved generalization. It introduces a tractable continuous‑time mean–fluctuation model that couples the window‑averaged trajectory with its fluctuation covariance, derives this model rigorously from a sharp‑valley framework, and analyzes its stationary states and linear stability. The authors also extend the model to wide two‑layer networks, deriving a Wasserstein‑2 gradient flow for weights and fluctuations, proving well‑posedness, a mean‑field limit, and conditional convergence results, with numerical experiments illustrating the predictions and finite‑time limitations.

arXiv Machine Learning
Jul 8

A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks

arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.

By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv Machine Learning
Sep 18

Learning-Induced Dynamical Transition in Recurrent Neural Networks

The paper presents a non-equilibrium dynamical mean-field theory (DMFT) that explains how learning reshapes the dynamics of recurrent neural networks, turning initially chaotic activity into stable, task-dependent behavior. It shows that a slow, feedback-driven learning process gradually increases effective feedback strength, driving the network through a bifurcation that marks the transition from chaotic to stable dynamics. By deriving the two-time correlation function, the authors identify a critical feedback strength and a learning-rate-dependent critical time that separate these regimes, and they demonstrate that the theory accurately predicts the network’s output evolution during training, matching numerical simulations.

By Varun Vaidya