arXiv Machine Learning

Learning to optimize with guarantees: a complete characterization of linearly convergent algorithms

arXiv:2508. 00775v2 Announce Type: replace-cross Abstract: The design of many classical optimization algorithms is driven by the certification of linear convergence rates over classes of optimization problems.

arXiv Machine Learning
Sep 14

High-Probability Convergence of SGD via Batched Updates

The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.

By Feng Zhu, Robert W. Heath Jr., Aritra Mitra
arXiv Machine Learning
1d ago

HUANet: Hard-Constrained Unrolled ADMM for Constrained Convex Optimization

HUANet is a deep neural network architecture that unrolls the Alternating Direction Method of Multipliers (ADMM) into a trainable model for accelerating parametric constrained convex optimization. It embeds a hard‑constrained neural network in each ADMM iteration, using a differentiable correction stage to enforce affine equalities of the primal subproblem. The method also incorporates first‑order optimality conditions into a self‑supervised training loss, and numerical experiments on benchmark problems and a control application demonstrate its effectiveness in speeding up constrained convex optimization.

By Trinh Tran, Binh Nguyen, Truong X. Nghiem
arXiv Machine Learning
4d ago

Learning Distributionally Robust First-Order Methods for Convex Optimization

The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.

By Vinit Ranjan, Jisun Park, Bartolomeo Stellato