arXiv:2603. 09923v4 Announce Type: replace Abstract: Exponential moving averages (EMAs) are a central component of widely used adaptive optimizers such as Adam.
By Ganzhao Yuan
The paper presents a theoretical study of Adam in non‑stationary stochastic optimization, distinguishing two regimes: Euclidean tracking under adaptive strong monotonicity and high‑probability projected stationarity for general smooth objectives. It derives finite‑time bounds that decompose into initialization, objective drift, first‑moment tracking error (β₁), and preconditioner perturbation (β₂), and characterizes burn‑in times for constant and step‑decay schedules. The analysis reveals a noise–drift tradeoff, showing that in noise‑dominated settings Adam’s adaptive mechanisms can improve guarantees, while in drift‑dominated settings they may worsen tracking, potentially making vanilla SGD preferable.
By Sharan Sahu, Abir Sarkar, Cameron J. Hogan, Martin T. Wells
arXiv:2607. 27383v1 Announce Type: new Abstract: We establish the first convergence guarantees for the plain vector-form \emph{Adam} optimizer under heavy-tailed stochastic noise.
By Yijiang Pang
The paper establishes uniform a priori bounds for the Adam optimizer, enabling an unconditional error analysis for a broad class of strongly convex stochastic optimization problems. Prior analyses were conditional, assuming Adam remained bounded, whereas this work removes that assumption. The results provide a rigorous foundation for Adam’s performance in training deep neural networks and other convex optimization tasks.
By Steffen Dereich, Thang Do, Arnulf Jentzen
arXiv:2608.30382v1 Announce Type: new
Abstract: Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW o...
By Steffen Dereich, Arnulf Jentzen
arXiv:2609.37787v1 Announce Type: new
Abstract: Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framew...
By Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou, Youzhi Zhang
arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).
By Steffen Dereich, Arnulf Jentzen, Adrian Riekert
arXiv:2608. 04026v1 Announce Type: cross Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization.
By Oluwasegun A. Somefun
arXiv:2609.36600v1 Announce Type: cross
Abstract: Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that emplo...
By Tao Jiang, Lin Xiao
arXiv:2605. 29547v2 Announce Type: replace-cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators.
By Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
By Yushun Zhang
arXiv:2505. 13196v3 Announce Type: replace-cross Abstract: We introduce Velocity-Regularized Adam (VRAdam), a physics-inspired optimizer for training deep neural networks that draws on ideas from quartic terms for kinetic energy with its stabilizing effects on various system dynamics.
By Pranav Vaidhyanathan, Lucas Schorling, Natalia Ares, Maike Osborne