arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
By Andrea Montanari, Kangjie Zhou
The paper introduces a learning-based surrogate approach for stochastic optimization problems where uncertainty depends on the decision, modeled via a nonparametric regression. It constructs a surrogate that embeds iteratively updated Jacobian estimates, using an adaptive random design that focuses sampling near the current iterate to achieve dimension‑independent convergence of the Jacobian estimates. The resulting learning‑based stochastic prox‑linear (L‑SPL) algorithm demonstrates nonasymptotic convergence rates and outperforms existing methods in sample efficiency and objective value in numerical experiments.
By Boyang Shen, Junyi Liu
arXiv:2608.30382v1 Announce Type: new
Abstract: Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW o...
By Steffen Dereich, Arnulf Jentzen
arXiv:2602. 02877v2 Announce Type: replace Abstract: This paper studies optimization for a family of problems termed $\textbf{compositional entropic risk minimization}$, in which each data's loss is formulated as a Log-Expectation-Exponential (Log-E-Exp) function.
By Xiyuan Wei, Linli Zhou, Bokun Wang, Chih-Jen Lin, Tianbao Yang
arXiv:2202. 08832v3 Announce Type: replace-cross Abstract: We study a general class of optimization problems with decision variable $\boldsymbol{\Theta} \in \mathbb{R}^{p \times k}$ and cost function which is the sum of $n$ terms, each dependent on $\boldsymbol{\Theta}$ through the $k$-dimensional projection $\boldsymbol{\Theta}^\top \boldsymbol{x}_i$, where $\boldsymbol{x}_i$, $i \leq n$ are i.
By Andrea Montanari, Basil Saeed
arXiv:2608. 05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function.
By Felipe Atenas, Alejandro Jofr\'e, Pedro P\'erez-Aros, David Torregrosa-Bel\'en
arXiv:2606. 14640v1 Announce Type: new Abstract: We study Online Convex Optimization (OCO) over a convex set $K\subseteq \mathbb R^d$, where in each round $t$ the learner selects $x_t\in K$ and then observes a convex loss $f_t:K\to[0,1]$, with the goal of minimizing regret to the best fixed decision in hindsight.
By Simone Di Gregorio, Anupam Gupta, Stefano Leonardi, Matteo Russo
arXiv:2604. 13130v2 Announce Type: replace Abstract: We study learning to learn through the lens of hyperparameter tuning.
By Saumya Goyal, Rohith Rongali, Ritabrata Ray, Barnab\'as P\'oczos
arXiv:2601. 18115v2 Announce Type: replace Abstract: We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of covariate distributions.
By Guyang Cao, Shuyao Li, Sushrut Karmalkar, Jelena Diakonikolas
arXiv:2606. 27298v1 Announce Type: cross Abstract: We study the fundamental problem of learning a high-dimensional Gaussian truncated to an unknown halfspace.
By Haitong Liu, Deepak Narayanan Sridharan, David Steurer, Manuel Wiedmer
In this work, we study the oracle complexity of finding an $ε$-stationary point for nonconvex-strongly-convex (NC-SC) bilevel optimization using only first-order oracles. Existing methods achieving th...
arXiv:2609. 20687v1 Announce Type: cross Abstract: We study first-order black-box convex optimization over an $\ell_p$-ball for objectives Lipschitz in the $\ell_q$-norm, solving in the affirmative the nonsmooth version of the COLT open question (Guz15b) on whether the geometry of a smaller feasible set ($p < q$) can improve convergence rates in convex optimization, and matching prior lower bounds up to logarithmic factors.
By David Mart\'inez-Rubio, Brian Bullins, Crist\'obal Guzm\'an, Mathieu Molina