arXiv:2606. 18080v1 Announce Type: new Abstract: Gradient descent in deep learning may operate at the edge of stability (EoS), a regime in which the largest eigenvalue of the loss Hessian hovers near the stability threshold $2/\eta$, where $\eta$ is the learning rate.
By Pierre Marion
arXiv:2301. 06308v2 Announce Type: replace-cross Abstract: Sharpness-aware minimization (SAM) is a training method that seeks to find flat minima in deep learning, resulting in state-of-the-art performance across various domains.
By Hoki Kim, Jinseong Park, Yujin Choi, Jaewook Lee
The paper investigates the "edge of stability" phenomenon in deep learning, where Hessian eigenvalues remain stable above a classically predicted unstable threshold. It shows that many first‑order optimizers, including gradient descent, can violate this stability bound by up to a factor of 21.1, and that this deviation depends systematically on the optimizer used. The authors propose a new stability threshold based on the directional Hessian and gradient‑alignment score, which removes optimizer‑dependent offsets and offers consistent predictions while providing diagnostic tools to understand how optimizers balance temporal and spatial budgets.
By Jaerin Lee, Kyoung Mu Lee
arXiv:2606. 30930v1 Announce Type: cross Abstract: Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory.
By Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun, Rene Vidal
arXiv:2608.20638v2 Announce Type: replace-cross
Abstract: The edge-of-stability (EoS) phenomenon of full-batch Adam has been widely observed, yet its underlying dynamical mechanism remains poorly und...
By Yiman Fong, Heng Yang
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic
The paper investigates how gradient descent behaves near codimension‑one bifurcations in recurrent neural networks by analyzing the global empirical Neural Tangent Kernel (GeNTK). Under local center‑manifold conditions, the parameter‑to‑state Jacobian is approximated by a low‑rank normal‑form operator, causing the GeNTK and Fisher information matrix to become strongly amplified and anisotropic, concentrating on a rank‑one or rank‑two channel depending on the bifurcation type. Experiments on high‑dimensional RNNs confirm that this low‑rank concentration coincides with abrupt loss changes, subtask interference, and aligns with changes in memory dynamics in a 15‑task LeakyRNN.
By James Hazelden, Eric Shea-Brown
arXiv:2609.01034v1 Announce Type: new
Abstract: The central flow of Cohen et al. (2025) is an empirically accurate continuous-time model of gradient descent at the edge of stability in deep learning,...
By Rapha\"el Berthier
The paper studies how Adam’s two momentum timescales, β1 and β3, influence loss spikes during neural‑network training. By mapping training dynamics across the (β1,β3) plane, the authors find an approximately linear boundary, 1-β3 = C(1-β1), that separates spiky from non‑spiky behavior, with the coefficient C linked to the effective loss exponent in superquadratic loss functions. They also show that confident cross‑entropy losses create a core–wall landscape that behaves superquadratically at the scale of an optimizer update, explaining the observed spikes.
By Gaoxiang Tang, Huanran Chen, Ziming Liu
arXiv:2510.25060v2 Announce Type: replace-cross
Abstract: In this work, we study the nonlinear dynamics of a shallow neural network trained with mean-squared loss and leaky ReLU activation. Under Gau...
By Jingzhou Liu
arXiv:2603. 05002v3 Announce Type: replace Abstract: The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian approaches and then hovers near the stability threshold $2/\eta$ during gradient descent (GD) with step size $\eta$.
By Rustem Islamov, Michael Crawshaw, Jeremy Cohen, Robert Gower
arXiv:2607. 13631v1 Announce Type: new Abstract: The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of generalization, second-order learning algorithms, etc.
By Jasraj Singh, Enea Monzio Compagnoni, Antonio Orvieto