Difference of Convex Programming in the Wasserstein Space with Applications to MMD Optimization
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
arXiv:2411. 15067v2 Announce Type: replace-cross Abstract: We investigate proximal descent methods, inspired by the minimizing movement scheme introduced by Jordan, Kinderlehrer and Otto, for optimizing entropy-regularized functionals on the Wasserstein space.
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
arXiv:2311. 15365v3 Announce Type: replace Abstract: We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space.
arXiv:2605. 30253v2 Announce Type: replace-cross Abstract: We study the contraction in Wasserstein distance of the coordinate ascent variational inference algorithm.
arXiv:2505. 07124v3 Announce Type: replace Abstract: We study inverse problems where an unknown potential is observed only through samples from the measure it induces by a convex variational principle.
arXiv:2607. 16384v1 Announce Type: new Abstract: For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons.
arXiv:2607. 02003v1 Announce Type: cross Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics.
arXiv:2605. 18694v2 Announce Type: replace-cross Abstract: Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process.
arXiv:2602. 04272v2 Announce Type: replace-cross Abstract: The Importance-Weighted Evidence Lower Bound (IW-ELBO) has emerged as an effective objective for variational inference (VI), tightening the standard ELBO and mitigating the mode-seeking behaviour.
arXiv:2409. 08469v4 Announce Type: replace-cross Abstract: We provide finite-particle convergence rates for the Stein Variational Gradient Descent (SVGD) algorithm in the Kernelized Stein Discrepancy ($\mathsf{KSD}$) and Wasserstein-2 metrics.
arXiv:2608. 05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function.
arXiv:2607. 07468v1 Announce Type: cross Abstract: We study the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning.