arXiv Machine Learning

A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization

arXiv:2505. 01258v2 Announce Type: replace-cross Abstract: Bilevel optimization has recently attracted significant attention in machine learning due to its wide range of applications and advanced hierarchical optimization capabilities.

arXiv Machine Learning
Jul 24

Non-Stationary Functional Bilevel Optimization

arXiv:2601. 15363v2 Announce Type: replace-cross Abstract: Functional bilevel optimization (FBO) provides a powerful framework for hierarchical learning in function spaces, yet current methods are limited to static offline settings and perform suboptimally in online, non-stationary scenarios.

By Jason Bohne, Ieva Petrulionyte, Michael Arbel, Julien Mairal, Pawe{\l} Polak
arXiv Statistics ML
Sep 11

Learning-Based Surrogate Method for Stochastic Optimization under Decision-Dependent Uncertainty with Adaptive Random Designs

The paper introduces a learning-based surrogate approach for stochastic optimization problems where uncertainty depends on the decision, modeled via a nonparametric regression. It constructs a surrogate that embeds iteratively updated Jacobian estimates, using an adaptive random design that focuses sampling near the current iterate to achieve dimension‑independent convergence of the Jacobian estimates. The resulting learning‑based stochastic prox‑linear (L‑SPL) algorithm demonstrates nonasymptotic convergence rates and outperforms existing methods in sample efficiency and objective value in numerical experiments.

By Boyang Shen, Junyi Liu
arXiv Machine Learning
1d ago

Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization

The paper studies stochastic multi‑level optimization where the objective is a nested composition of smooth non‑convex functions. It introduces momentum‑based estimators that track function values at each level, achieving an optimal sample complexity of ≠(ε⁻⁴) for finding an ε‑stationary point without relying on average smoothness assumptions. The authors also present a batch‑free variant using first‑order approximations and clipping, and demonstrate the methods on risk‑averse portfolio optimization and hierarchical tilted empirical risk minimization.

By Wei Jiang, Rui Yan, Sifan Yang, Yuanyu Wan, Lijun Zhang, Zechao Li