arXiv Machine Learning

Stochastic Heavy Ball with Polyak Step Size and Armijo Line Search: A General Convergence Analysis

arXiv Machine Learning
Aug 10

A proximal subgradient method for nonconvex stochastic optimization under the Kurdyka-{\L}ojasiewicz condition

arXiv:2608. 05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function.

By Felipe Atenas, Alejandro Jofr\'e, Pedro P\'erez-Aros, David Torregrosa-Bel\'en
arXiv Machine Learning
Sep 14

Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

The paper proves that stochastic gradient descent with gradient clipping and additive Gaussian noise (SGD‑CN) converges almost surely under smoothness and bounded noise assumptions, given standard decaying step sizes. The analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, showing that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods in both convex and nonconvex regimes.

By Amartya Mukherjee, Jun Liu
arXiv Machine Learning
Sep 23

Polyak-Type Extragradient Methods for Monotone Root-Finding Problems

The paper investigates Polyak-type step-size strategies for extragradient methods applied to deterministic and stochastic monotone root-finding problems. It shows that the projection-based correction in deterministic extragradient can be derived by minimizing an upper bound on the distance to a solution, mirroring classical Polyak step-size construction. The authors provide a unified deterministic analysis that does not require global Lipschitz continuity, achieving sublinear convergence under H"older or “(L0, L1)-Lipschitz” conditions and linear convergence with strong monotonicity, and extend the approach to stochastic settings with both direct and decreasing step-size variants.

By TaeHo Yoon, Sayantan Choudhury, Ezra Greenberg, Nicolas Loizou
arXiv Machine Learning
Jul 2

Towards Weaker Variance Assumptions for Stochastic Optimization

arXiv:2504. 09951v2 Announce Type: replace-cross Abstract: We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable.

By Ahmet Alacaoglu, Yura Malitsky, Stephen J. Wright
arXiv Machine Learning
Sep 21

Single-Loop Stochastic Projected Damped Extragradient Methods for Stochastic Nonconvex--(Strongly) Concave Minimax Optimization

The paper introduces single-loop stochastic projected damped extragradient (SPDE) and its variance-reduced variant (VR-SPDE) for stochastic nonconvex–(strongly) concave minimax problems. It provides SFO complexity bounds for achieving game stationarity and optimization stationarity, improving upon previous multi-loop methods while maintaining a single-loop structure. The results claim the best-known SFO complexities for these stationarity criteria among single-loop stochastic first‑order methods.

By Huiling Zhang, Minhao Zhang, Zi Xu