A new class of asynchronous adaptive first-order optimization methods is introduced, comprising asynchronous variants of several popular algorithms. Versions of these methods using momentum and/or inexact normalization are also considered.
arXiv:2509. 14969v2 Announce Type: replace Abstract: We introduce a new adaptive step-size strategy for convex optimization with stochastic gradient that exploits the local geometry of the objective function only by means of a first-order stochastic oracle and without any hyper-parameter tuning.
By Jean-Fran\c{c}ois Aujol, J\'er\'emie Bigot, Camille Castera
The paper introduces a unified framework for first‑order optimization algorithms applied to nonconvex unconstrained problems. It incorporates adaptively preconditioned gradients and covers popular methods such as full and diagonal AdaGrad, AdaNorm, and an adaptive variant of Muon. The framework supports heterogeneous geometries across variable groups and provides a fully stochastic global convergence analysis for all methods, with or without two types of momentum, under reasonable variance assumptions without requiring bounded stochastic gradients or small step sizes.
By S. Gratton, Ph. L. Toint
arXiv:2605. 18694v2 Announce Type: replace-cross Abstract: Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process.
By Zijian Liu
arXiv:2502.21099v3 Announce Type: replace-cross
Abstract: This paper proposes {\sf AEPG-SPIDER}, an Adaptive Extrapolated Proximal Gradient (AEPG) method with variance reduction for minimizing compos...
By Ganzhao Yuan
The paper examines the challenge of bridging the performance gap between homogeneous and heterogeneous asynchronous optimization in large-scale machine learning. It demonstrates that under common first- and second-order similarity assumptions, no randomized algorithm can improve the pessimistic time complexity bounds for heterogeneous settings. The authors further show that even weak interpolation is insufficient, but by combining strong interpolation with a local Polyak‑Lojasiewicz condition, they achieve a new time complexity that matches the best-known homogeneous result without requiring identical data distributions.
By Alexander Tyurin
arXiv:2607. 28902v1 Announce Type: new Abstract: We develop a parallel framework that assembles static gradient methods to achieve better adaptivity.
By Bin Fu
arXiv:2603. 09923v4 Announce Type: replace Abstract: Exponential moving averages (EMAs) are a central component of widely used adaptive optimizers such as Adam.
By Ganzhao Yuan
arXiv:2606. 00520v1 Announce Type: cross Abstract: Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite $p$-th moment for $p\in\left(1,2\right)$, a setting known as the heavy-tailed noise assumption.
By Zijian Liu
The paper introduces a parallel architecture for stochastic gradient methods that adaptively selects the number of iterations. An algorithm A(x₀, y) takes an initial point and a step limit y, and p processors search for an appropriate iteration count T using a prescribed function h. The framework guarantees a (p, αₚ)-approximation, meaning for any T ≥ T₀ there exists a processor and stage where the cumulative iterations lie within a factor αₚ of T, and the authors prove tight lower bounds for αₚ while presenting simple arithmetic stochastic gradient methods that use only divisions by powers of two.
By Bin Fu
arXiv:2606. 08783v1 Announce Type: cross Abstract: Orthogonalized momentum updates, as used in Muon-style optimizers, have recently shown strong empirical stability in large-scale deep learning.
By Ganzhao Yuan
arXiv:2607. 26562v1 Announce Type: cross Abstract: We study optimization under performative prediction, where deploying a model affects the future data distribution.
By Hiroki Hamaguchi, Yuya Hikima, Hiroshi Sawada, Akiko Takeda