arXiv AI

Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients

arXiv Machine Learning
Sep 14

Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization

The paper presents a theoretical study of Adam in non‑stationary stochastic optimization, distinguishing two regimes: Euclidean tracking under adaptive strong monotonicity and high‑probability projected stationarity for general smooth objectives. It derives finite‑time bounds that decompose into initialization, objective drift, first‑moment tracking error (β₁), and preconditioner perturbation (β₂), and characterizes burn‑in times for constant and step‑decay schedules. The analysis reveals a noise–drift tradeoff, showing that in noise‑dominated settings Adam’s adaptive mechanisms can improve guarantees, while in drift‑dominated settings they may worsen tracking, potentially making vanilla SGD preferable.

By Sharan Sahu, Abir Sarkar, Cameron J. Hogan, Martin T. Wells
arXiv Machine Learning
Sep 14

Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

The paper proves that stochastic gradient descent with gradient clipping and additive Gaussian noise (SGD‑CN) converges almost surely under smoothness and bounded noise assumptions, given standard decaying step sizes. The analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, showing that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods in both convex and nonconvex regimes.

By Amartya Mukherjee, Jun Liu
arXiv Machine Learning
Sep 2

Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method

The paper establishes uniform a priori bounds for the Adam optimizer, enabling an unconditional error analysis for a broad class of strongly convex stochastic optimization problems. Prior analyses were conditional, assuming Adam remained bounded, whereas this work removes that assumption. The results provide a rigorous foundation for Adam’s performance in training deep neural networks and other convex optimization tasks.

By Steffen Dereich, Thang Do, Arnulf Jentzen