arXiv Machine Learning By Vihaan Paka-Hegde

How Far is Adam from Natural Gradient Descent?

Read the original on arXiv Machine Learning →

The paper investigates how Adam’s update rule relates to natural gradient descent (NGD) by treating Adam as a diagonal empirical Fisher approximation with additional factors such as diagonal truncation, empirical label substitution, and temporal lag. Using a scale‑invariant metric, the authors quantify Adam’s geometric deviation from true NGD across four loss landscapes—well‑conditioned and ill‑conditioned linear regression, logistic regression, and a small neural network—finding that deviation is low in well‑conditioned settings but can reach about 10³ in ill‑conditioned or non‑convex scenarios. Despite higher geometric drift correlating with slower early optimization, Adam still achieves low final loss, and the improved empirical Fisher (iEF) yields more stable trajectories than the standard empirical Fisher (EF).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

Why $\beta_1 = \beta_2$ Is Dynamically Special in Adam

The paper investigates why setting the two momentum parameters of Adam equal (β1=β2) has a special dynamic effect. By analysing Adam in continuous time, the authors show that the update decomposes into a sign component, a magnitude‑lag term proportional to the difference between the two memory times, and other terms. This lag term disappears exactly when β1=β2, making the diagonal the only regime where the mismatch‑induced response is structurally absent. Experiments on six vision and language tasks confirm that tied configurations are sign‑dominated, have smaller lag contributions, and exhibit smoother update‑norm trajectories.

By Alberto Fern\'andez-Hern\'andez, Cristian P\'erez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ort\'i
arXiv Machine Learning
Sep 17

Beyond Quadratic Loss: The Stability Phase Diagram of Adam

The paper studies how Adam’s two momentum timescales, β1 and β3, influence loss spikes during neural‑network training. By mapping training dynamics across the (β1,β3) plane, the authors find an approximately linear boundary, 1-β3 = C(1-β1), that separates spiky from non‑spiky behavior, with the coefficient C linked to the effective loss exponent in superquadratic loss functions. They also show that confident cross‑entropy losses create a core–wall landscape that behaves superquadratically at the scale of an optimizer update, explaining the observed spikes.

By Gaoxiang Tang, Huanran Chen, Ziming Liu