arXiv Machine Learning

Sharp First-Order Lower Bounds for Higher-Order Smooth Nonconvex Optimization

arXiv:2606. 05438v1 Announce Type: new Abstract: We study the deterministic first-order oracle complexity of finding \(\epsilon\)-stationary points in smooth nonconvex optimization when the objective satisfies higher-order smoothness assumptions.

Hugging Face Trending Papers
Sep 24

Anchored Extra-Proximal Methods: Optimal Higher-Order Methods for Monotone Inclusion Problems

The paper introduces the Anchored Extra-Proximal (AEP) framework for solving composite monotone inclusion problems, combining anchored extrapolation with an inexact anchored proximal update. By replacing the operator in the implicit update with its Taylor approximation and using a bisection line search, the authors derive a pth-order method that achieves a tangent-residual error ε in “~O(ε^{-2/(3p-1)})” oracle calls for every p ≥ 2. This complexity matches a proven lower bound, establishing the method as optimally efficient for deterministic algorithms in the pth-order oracle model.

arXiv Machine Learning
Sep 3

Improved Gradient Descent Lower Bounds Beyond Nesterov

The paper investigates the limits of accelerating gradient descent (GD) using predetermined step sizes in smooth convex optimization. It establishes new lower bounds: an ≥·n−1.6342 non‑anytime bound and an ≥·n−1.2408 anytime bound, surpassing previous results. These findings also demonstrate a strict separation between convergence exponents achievable in non‑anytime versus anytime settings.

By Yuhan Ye, Kaizhao Liu
arXiv Machine Learning
1d ago

Optimal Stochastic Bilevel Optimization with First-Order Oracles

The paper investigates nonconvex–strongly-convex bilevel optimization using a stochastic first-order oracle. It introduces MRT‑FD, a single-loop first‑order algorithm that tracks the upper-level variable, the lower-level solution, and an auxiliary response from implicit differentiation, updating all variables in each iteration and approximating second‑order derivative actions via order‑p finite differences. For any fixed finite smoothness order p ≥ 1, MRT‑FD achieves an ε‑stationary point with O(ε^{‑4‑2/p}) stochastic gradient queries, and the authors prove a matching Ω(ε^{‑4‑2/p}) lower bound, thereby closing the complexity gap in this setting.

By Linxuan Pan, Junchi Yang