arXiv:2607. 16261v2 Announce Type: replace-cross Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments.
By Apostolos Avranas
The paper introduces a learning-based surrogate approach for stochastic optimization problems where uncertainty depends on the decision, modeled via a nonparametric regression. It constructs a surrogate that embeds iteratively updated Jacobian estimates, using an adaptive random design that focuses sampling near the current iterate to achieve dimension‑independent convergence of the Jacobian estimates. The resulting learning‑based stochastic prox‑linear (L‑SPL) algorithm demonstrates nonasymptotic convergence rates and outperforms existing methods in sample efficiency and objective value in numerical experiments.
By Boyang Shen, Junyi Liu
The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.
By Feng Zhu, Robert W. Heath Jr., Aritra Mitra
arXiv:2606. 01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy.
By Wyame Benslimane, Tinghan Ye, Pascal Van Hentenryck, Paul Grigas
arXiv:2609.36310v1 Announce Type: new
Abstract: Everywhere learning provides a principled framework for training AI models under constraints that must hold throughout the data distribution. In the du...
By Ignacio Boero, Jonathan Nixon, Alejandro Ribeiro
arXiv:2607. 16261v1 Announce Type: cross Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments.
By Apostolos Avranas
arXiv:2307. 05213v3 Announce Type: replace-cross Abstract: Many real-world optimization problems contain parameters that are unknown before deployment time, either due to stochasticity or to lack of information (e.
By Mattia Silvestri, Senne Berden, Jayanta Mandi, Ali \.Irfan Mahmuto\u{g}ullar{\i}, Brandon Amos, Tias Guns, Michele Lombardi
arXiv:2505.05261v4 Announce Type: replace-cross
Abstract: Two-stage stochastic programming (2SP) offers a basic framework for modelling decision-making under uncertainty, yet scalability remains a ch...
By Yu Liu, Fabricio Oliveira, Jan Kronqvist
arXiv:2607. 24007v1 Announce Type: cross Abstract: We revisit contextual optimization from the perspective of policy class design.
By Zikun Lin, Rui Chen, Yijie Wang
arXiv:2502. 05349v2 Announce Type: replace-cross Abstract: Two-stage stochastic programs (2SPs) are widely used for decision-making under uncertainty, but their practical deployment is often limited by the large number of scenarios needed to approximate the conditional distribution of uncertain outcomes.
By David Islip, Roy H. Kwon, Sanghyeon Bae, Woo Chang Kim
The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.
By Vinit Ranjan, Jisun Park, Bartolomeo Stellato
arXiv:2608. 09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success.
By Binchuan Qi