Stochastic Heavy Ball with Polyak Step Size and Armijo Line Search: A General Convergence Analysis
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 00520v1 Announce Type: cross Abstract: Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite $p$-th moment for $p\in\left(1,2\right)$, a setting known as the heavy-tailed noise assumption.
arXiv:2608. 05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function.
The paper proves that stochastic gradient descent with gradient clipping and additive Gaussian noise (SGD‑CN) converges almost surely under smoothness and bounded noise assumptions, given standard decaying step sizes. The analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, showing that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods in both convex and nonconvex regimes.
The paper investigates Polyak-type step-size strategies for extragradient methods applied to deterministic and stochastic monotone root-finding problems. It shows that the projection-based correction in deterministic extragradient can be derived by minimizing an upper bound on the distance to a solution, mirroring classical Polyak step-size construction. The authors provide a unified deterministic analysis that does not require global Lipschitz continuity, achieving sublinear convergence under H"older or “(L0, L1)-Lipschitz” conditions and linear convergence with strong monotonicity, and extend the approach to stochastic settings with both direct and decreasing step-size variants.
arXiv:2506.04192v4 Announce Type: replace-cross Abstract: Stochastic Frank-Wolfe is a classical optimization method for solving constrained optimization problems. On the other hand, recent optimizers...
arXiv:2605.28517v2 Announce Type: replace-cross Abstract: Stochastic gradient descent with momentum (SGDM) is one of the most widely used optimization algorithms in machine learning. While optimizati...