Difference of Convex Programming in the Wasserstein Space with Applications to MMD Optimization
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
arXiv:2411. 15067v2 Announce Type: replace-cross Abstract: We investigate proximal descent methods, inspired by the minimizing movement scheme introduced by Jordan, Kinderlehrer and Otto, for optimizing entropy-regularized functionals on the Wasserstein space.
arXiv:2311. 15365v3 Announce Type: replace Abstract: We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space.
The paper introduces a geometric framework for reinforcement learning that treats policies as mappings into the Wasserstein space of action probabilities. It establishes a Riemannian structure induced by stationary distributions, defines the tangent space of policies, and characterizes geodesics while addressing measurability concerns. The authors formulate a general RL optimization problem, construct a gradient flow via Otto's calculus, compute the gradient and Hessian of the energy, and demonstrate the approach with numerical examples for low‑dimensional problems and neural‑network‑parameterized policies for high‑dimensional settings.
arXiv:2102. 09235v3 Announce Type: replace Abstract: Recent studies revealed the mathematical connection between deep neural networks (DNNs) and dynamic systems.
The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.
arXiv:2502. 17602v2 Announce Type: replace-cross Abstract: We study a class of stochastic nonsmooth optimization problems in which an outer variable minimizes the expectation of a pointwise maximum.
arXiv:2602. 04272v2 Announce Type: replace-cross Abstract: The Importance-Weighted Evidence Lower Bound (IW-ELBO) has emerged as an effective objective for variational inference (VI), tightening the standard ELBO and mitigating the mode-seeking behaviour.
The paper introduces a new convergence framework for solving distributionally robust optimization problems formulated as nonconvex, nonconcave minimax problems over a Euclidean space and a Riemannian manifold. It defines a "basin saddle point"—a locally defined Nash equilibrium—and proves that a Riemannian gradient ascent–descent algorithm converges to such points under a local Łojasiewicz growth condition. The authors apply this theory to a statistical risk DRO problem over Gaussian measures, deriving explicit convergence rates and constants in terms of data dimension, loss moments, and reference covariance.
arXiv:2607. 02003v1 Announce Type: cross Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics.
arXiv:2506. 04480v2 Announce Type: replace-cross Abstract: This paper focuses on Geodesic Principal Component Analysis (GPCA) on a collection of probability distributions using the Otto-Wasserstein geometry.
This paper introduces a generative model that minimizes the second‑order Wasserstein loss (W₂) by solving a distribution‑dependent ordinary differential equation (ODE) whose dynamics involve the Kantorovich potential of the true data distribution and its current estimate. The authors prove that the time‑marginal laws of this ODE form a gradient flow for the W₂ loss, converging exponentially to the true data distribution, and propose an Euler scheme that recovers this gradient flow in the limit. An algorithm based on this scheme, combined with persistent training, is shown in experiments to outperform Wasserstein GANs in both low‑ and high‑dimensional settings when the level of persistent training is appropriately increased.