arXiv Machine Learning

Stochastic Nonconvex Bilevel Optimization: Improved Rates Without Rare-Visit Assumption

arXiv:2609. 06580v1 Announce Type: cross Abstract: We investigate stochastic simple bilevel optimization with smooth and possibly nonconvex upper- and lower-level objectives.

arXiv Machine Learning
1d ago

Optimal Stochastic Bilevel Optimization with First-Order Oracles

The paper investigates nonconvex–strongly-convex bilevel optimization using a stochastic first-order oracle. It introduces MRT‑FD, a single-loop first‑order algorithm that tracks the upper-level variable, the lower-level solution, and an auxiliary response from implicit differentiation, updating all variables in each iteration and approximating second‑order derivative actions via order‑p finite differences. For any fixed finite smoothness order p ≥ 1, MRT‑FD achieves an ε‑stationary point with O(ε^{‑4‑2/p}) stochastic gradient queries, and the authors prove a matching Ω(ε^{‑4‑2/p}) lower bound, thereby closing the complexity gap in this setting.

By Linxuan Pan, Junchi Yang
arXiv Machine Learning
Sep 21

Single-Loop Stochastic Projected Damped Extragradient Methods for Stochastic Nonconvex--(Strongly) Concave Minimax Optimization

The paper introduces single-loop stochastic projected damped extragradient (SPDE) and its variance-reduced variant (VR-SPDE) for stochastic nonconvex–(strongly) concave minimax problems. It provides SFO complexity bounds for achieving game stationarity and optimization stationarity, improving upon previous multi-loop methods while maintaining a single-loop structure. The results claim the best-known SFO complexities for these stationarity criteria among single-loop stochastic first‑order methods.

By Huiling Zhang, Minhao Zhang, Zi Xu
arXiv Machine Learning
Jul 2

Towards Weaker Variance Assumptions for Stochastic Optimization

arXiv:2504. 09951v2 Announce Type: replace-cross Abstract: We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable.

By Ahmet Alacaoglu, Yura Malitsky, Stephen J. Wright
arXiv Machine Learning
5d ago

To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity

arXiv:2609. 30501v1 Announce Type: new Abstract: Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.

By Zhiyao Zhang, Menglu Yu, Alvaro Velasquez, Nathaniel D. Bastian, Jia Liu