A new class of asynchronous adaptive first-order optimization methods is introduced, comprising asynchronous variants of several popular algorithms. Versions of these methods using momentum and/or inexact normalization are also considered.
arXiv:2509. 14969v2 Announce Type: replace Abstract: We introduce a new adaptive step-size strategy for convex optimization with stochastic gradient that exploits the local geometry of the objective function only by means of a first-order stochastic oracle and without any hyper-parameter tuning.
By Jean-Fran\c{c}ois Aujol, J\'er\'emie Bigot, Camille Castera
The paper introduces a unified framework for first‑order optimization algorithms applied to nonconvex unconstrained problems. It incorporates adaptively preconditioned gradients and covers popular methods such as full and diagonal AdaGrad, AdaNorm, and an adaptive variant of Muon. The framework supports heterogeneous geometries across variable groups and provides a fully stochastic global convergence analysis for all methods, with or without two types of momentum, under reasonable variance assumptions without requiring bounded stochastic gradients or small step sizes.
By S. Gratton, Ph. L. Toint
arXiv:2605. 18694v2 Announce Type: replace-cross Abstract: Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process.
By Zijian Liu
arXiv:2502.21099v3 Announce Type: replace-cross
Abstract: This paper proposes {\sf AEPG-SPIDER}, an Adaptive Extrapolated Proximal Gradient (AEPG) method with variance reduction for minimizing compos...
By Ganzhao Yuan
The paper examines the challenge of bridging the performance gap between homogeneous and heterogeneous asynchronous optimization in large-scale machine learning. It demonstrates that under common first- and second-order similarity assumptions, no randomized algorithm can improve the pessimistic time complexity bounds for heterogeneous settings. The authors further show that even weak interpolation is insufficient, but by combining strong interpolation with a local Polyak‑Lojasiewicz condition, they achieve a new time complexity that matches the best-known homogeneous result without requiring identical data distributions.
By Alexander Tyurin