Finding a stationary point of a stochastic convex problem
arXiv:2607. 06883v1 Announce Type: cross Abstract: We consider the problem of finding stationary points for stochastic convex optimization problems.
arXiv:2608. 05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function.
arXiv:2607. 06883v1 Announce Type: cross Abstract: We consider the problem of finding stationary points for stochastic convex optimization problems.
arXiv:2608. 06182v1 Announce Type: cross Abstract: We study stochastic extragradient (SEG) methods for solving monotone variational inequality problems (VIPs) over a feasible set.
The paper investigates Polyak-type step-size strategies for extragradient methods applied to deterministic and stochastic monotone root-finding problems. It shows that the projection-based correction in deterministic extragradient can be derived by minimizing an upper bound on the distance to a solution, mirroring classical Polyak step-size construction. The authors provide a unified deterministic analysis that does not require global Lipschitz continuity, achieving sublinear convergence under H"older or “(L0, L1)-Lipschitz” conditions and linear convergence with strong monotonicity, and extend the approach to stochastic settings with both direct and decreasing step-size variants.
arXiv:2504. 09951v2 Announce Type: replace-cross Abstract: We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable.
arXiv:2505.20817v3 Announce Type: replace-cross Abstract: Gradient clipping is widely used in language-model training to control heavy-tailed gradient noise and can improve convergence guarantees ove...
arXiv:2506.04192v4 Announce Type: replace-cross Abstract: Stochastic Frank-Wolfe is a classical optimization method for solving constrained optimization problems. On the other hand, recent optimizers...
arXiv:2502.21099v3 Announce Type: replace-cross Abstract: This paper proposes {\sf AEPG-SPIDER}, an Adaptive Extrapolated Proximal Gradient (AEPG) method with variance reduction for minimizing compos...
arXiv:2609.37425v1 Announce Type: cross Abstract: Fix a target accuracy $\varepsilon$, a gradient-noise level $s$, and a horizon $N$. We wish to design algorithms which minimize the probability of ob...
arXiv:2608. 03001v1 Announce Type: cross Abstract: Unit excitation (UE) is a common assumption in stochastic saddle avoidance: the stochastic error must have a uniformly positive component along every direction, in expectation.
arXiv:2605. 18694v2 Announce Type: replace-cross Abstract: Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process.
arXiv:2510.11676v2 Announce Type: replace-cross Abstract: We study convex composite optimization problems, where the objective function is given by the sum of a prox-friendly function and a convex fu...
In this work, we study the oracle complexity of finding an $ε$-stationary point for nonconvex-strongly-convex (NC-SC) bilevel optimization using only first-order oracles. Existing methods achieving th...