arXiv AI

Adam at the Edge of Stability: Adaptive Feedback, Provable Oscillation, and Gradient Reversal

arXiv Machine Learning
1d ago

How Far is Adam from Natural Gradient Descent?

The paper investigates how Adam’s update rule relates to natural gradient descent (NGD) by treating Adam as a diagonal empirical Fisher approximation with additional factors such as diagonal truncation, empirical label substitution, and temporal lag. Using a scale‑invariant metric, the authors quantify Adam’s geometric deviation from true NGD across four loss landscapes—well‑conditioned and ill‑conditioned linear regression, logistic regression, and a small neural network—finding that deviation is low in well‑conditioned settings but can reach about 10³ in ill‑conditioned or non‑convex scenarios. Despite higher geometric drift correlating with slower early optimization, Adam still achieves low final loss, and the improved empirical Fisher (iEF) yields more stable trajectories than the standard empirical Fisher (EF).

By Vihaan Paka-Hegde
arXiv Machine Learning
Jun 29

Adaptive Momentum and Nonlinear Damping for Neural Network Training

arXiv:2602. 00334v2 Announce Type: replace Abstract: Momentum Stochastic Gradient Descent (mSGD) relies on a fixed momentum coefficient shared across all parameters, failing to account for the heterogeneous structure of modern loss landscapes.

By Aikaterini Karoni, Rajit Rajpal, Benedict Leimkuhler, Gabriel Stoltz
arXiv AI
Sep 18

Why $\beta_1 = \beta_2$ Is Dynamically Special in Adam

The paper investigates why setting the two momentum parameters of Adam equal (β1=β2) has a special dynamic effect. By analysing Adam in continuous time, the authors show that the update decomposes into a sign component, a magnitude‑lag term proportional to the difference between the two memory times, and other terms. This lag term disappears exactly when β1=β2, making the diagonal the only regime where the mismatch‑induced response is structurally absent. Experiments on six vision and language tasks confirm that tied configurations are sign‑dominated, have smaller lag contributions, and exhibit smoother update‑norm trajectories.

By Alberto Fern\'andez-Hern\'andez, Cristian P\'erez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ort\'i
arXiv Machine Learning
Sep 17

Beyond Quadratic Loss: The Stability Phase Diagram of Adam

The paper studies how Adam’s two momentum timescales, β1 and β3, influence loss spikes during neural‑network training. By mapping training dynamics across the (β1,β3) plane, the authors find an approximately linear boundary, 1-β3 = C(1-β1), that separates spiky from non‑spiky behavior, with the coefficient C linked to the effective loss exponent in superquadratic loss functions. They also show that confident cross‑entropy losses create a core–wall landscape that behaves superquadratically at the scale of an optimizer update, explaining the observed spikes.

By Gaoxiang Tang, Huanran Chen, Ziming Liu