arXiv Machine Learning By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen

A Gradient Flow Perspective on Minimum MMD Estimation

Read the original on arXiv Machine Learning →

arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 16

Variance Reduction for Non-Log-Concave Sampling with Applications to Inverse Problems

arXiv:2606. 16257v1 Announce Type: cross Abstract: Sampling from high-dimensional, non-log-concave distributions with unnormalized densities is a fundamental challenge in machine learning, particularly when the exact gradient of the potential is unavailable and must be approximated via stochastic gradients that exhibit high variance under a fixed budget of gradient computations per iteration.

By M. Berk Sahin, Ahmet Ege Tanriverdi, Behzad Sharif, Abolfazl Hashemi
arXiv Machine Learning
1d ago

Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization

The paper studies stochastic multi‑level optimization where the objective is a nested composition of smooth non‑convex functions. It introduces momentum‑based estimators that track function values at each level, achieving an optimal sample complexity of ≠(ε⁻⁴) for finding an ε‑stationary point without relying on average smoothness assumptions. The authors also present a batch‑free variant using first‑order approximations and clipping, and demonstrate the methods on risk‑averse portfolio optimization and hierarchical tilted empirical risk minimization.

By Wei Jiang, Rui Yan, Sifan Yang, Yuanyu Wan, Lijun Zhang, Zechao Li
arXiv Machine Learning
Sep 14

High-Probability Convergence of SGD via Batched Updates

The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.

By Feng Zhu, Robert W. Heath Jr., Aritra Mitra