AdamX is a new first‑order optimizer that uses cosine similarity to adaptively control update magnitudes, making it scalable, model‑agnostic, and easy to add to existing training pipelines. It also includes a variance rectification scheme that smooths optimization early in training. Empirical results show AdamX achieves competitive convergence rates across various benchmark datasets and architectures, measured by the number of epochs needed to hit predefined performance thresholds under a fixed hyperparameter budget.
By Francisco Caldas, Ruben Belo, Cl\'audia Soares
The paper investigates how Adam’s update rule relates to natural gradient descent (NGD) by treating Adam as a diagonal empirical Fisher approximation with additional factors such as diagonal truncation, empirical label substitution, and temporal lag. Using a scale‑invariant metric, the authors quantify Adam’s geometric deviation from true NGD across four loss landscapes—well‑conditioned and ill‑conditioned linear regression, logistic regression, and a small neural network—finding that deviation is low in well‑conditioned settings but can reach about 10³ in ill‑conditioned or non‑convex scenarios. Despite higher geometric drift correlating with slower early optimization, Adam still achieves low final loss, and the improved empirical Fisher (iEF) yields more stable trajectories than the standard empirical Fisher (EF).
By Vihaan Paka-Hegde
arXiv:2602. 10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization.
By Francisco Patitucci, Aryan Mokhtari
arXiv:2608. 15824v1 Announce Type: new Abstract: Adam retains a moving average of past squared gradients in its denominator, but the optimization cost of this memory is not well understood.
By Jeonseong Kim
The paper establishes uniform a priori bounds for the Adam optimizer, enabling an unconditional error analysis for a broad class of strongly convex stochastic optimization problems. Prior analyses were conditional, assuming Adam remained bounded, whereas this work removes that assumption. The results provide a rigorous foundation for Adam’s performance in training deep neural networks and other convex optimization tasks.
By Steffen Dereich, Thang Do, Arnulf Jentzen
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
By Yushun Zhang