arXiv Machine Learning

A Gradient Flow Perspective on Minimum MMD Estimation

arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.

arXiv AI
Jun 16

Variance Reduction for Non-Log-Concave Sampling with Applications to Inverse Problems

arXiv:2606. 16257v1 Announce Type: cross Abstract: Sampling from high-dimensional, non-log-concave distributions with unnormalized densities is a fundamental challenge in machine learning, particularly when the exact gradient of the potential is unavailable and must be approximated via stochastic gradients that exhibit high variance under a fixed budget of gradient computations per iteration.

By M. Berk Sahin, Ahmet Ege Tanriverdi, Behzad Sharif, Abolfazl Hashemi
arXiv Machine Learning
1d ago

Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization

The paper studies stochastic multi‑level optimization where the objective is a nested composition of smooth non‑convex functions. It introduces momentum‑based estimators that track function values at each level, achieving an optimal sample complexity of ≠(ε⁻⁴) for finding an ε‑stationary point without relying on average smoothness assumptions. The authors also present a batch‑free variant using first‑order approximations and clipping, and demonstrate the methods on risk‑averse portfolio optimization and hierarchical tilted empirical risk minimization.

By Wei Jiang, Rui Yan, Sifan Yang, Yuanyu Wan, Lijun Zhang, Zechao Li
arXiv Machine Learning
Sep 14

High-Probability Convergence of SGD via Batched Updates

The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.

By Feng Zhu, Robert W. Heath Jr., Aritra Mitra
arXiv Machine Learning
4d ago

Learning Distributionally Robust First-Order Methods for Convex Optimization

The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.

By Vinit Ranjan, Jisun Park, Bartolomeo Stellato
arXiv Machine Learning
Sep 3

Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle

The paper revisits median‑of‑means estimation from a deterministic optimization perspective, introducing a family of block‑Lp estimators (for 0 < p ≤ 1) that achieve robust learning with heavy‑tailed and adversarially corrupted data. It shows that any convex block M‑estimator cannot attain the trimmed‑block oracle constant, while the nonconvex block‑Lp family provides finite‑sample robustness bounds that approach this oracle constant as p decreases. The authors also prove that the block‑Lp objectives have a benign landscape—every local minimum is close to the true parameter—and combine these results with block‑level concentration to obtain sub‑Gaussian deviation bounds under finite 2+δ moments, extending to high‑dimensional robust mean estimation and sparse regression.

By Angshul Majumdar
arXiv AI
Sep 10

Deep Barycentric Regression for Optimal Transport Map Estimation and its Statistical Optimality

The paper introduces BROT, a two‑step approach for estimating optimal transport maps. First, it computes the unregularized OT plan, then fits a deep neural network to the resulting barycentric targets using least‑squares regression. The authors prove that, under standard regularity conditions, BROT achieves the minimax convergence rate when the true OT map is Lipschitz, and demonstrate its effectiveness on synthetic data, images, and downstream tasks such as single‑cell perturbation prediction and unsupervised domain adaptation.

By Kunwoong Kim, Insung Kong, Yongdai Kim