arXiv:2501.14993v4 Announce Type: replace-cross
Abstract: The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent...
By Shuailong Zhu, Xiaohui Chen
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
By Cl\'ement Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2311. 15365v3 Announce Type: replace Abstract: We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space.
By Noboru Isobe
arXiv:2605. 30253v2 Announce Type: replace-cross Abstract: We study the contraction in Wasserstein distance of the coordinate ascent variational inference algorithm.
By Rocco Caprio, Adrien Corenflos, Sam Power
arXiv:2609. 27008v1 Announce Type: cross Abstract: We study the long-time behavior of Wasserstein gradient flows for interaction energies \[ \mathsf E[\mu] = \frac12\iint_{M\times M}K(x,y)\,\mathrm d\mu(x)\,\mathrm d\mu(y) \] on a closed manifold $M$.
By Zhengjiang Lin, Philippe Rigollet
The paper introduces a new convergence framework for solving distributionally robust optimization problems formulated as nonconvex, nonconcave minimax problems over a Euclidean space and a Riemannian manifold. It defines a "basin saddle point"—a locally defined Nash equilibrium—and proves that a Riemannian gradient ascent–descent algorithm converges to such points under a local Łojasiewicz growth condition. The authors apply this theory to a statistical risk DRO problem over Gaussian measures, deriving explicit convergence rates and constants in terms of data dimension, loss moments, and reference covariance.
By Rishabh Dixit, Pranav Upadrashta, Alex Cloninger
arXiv:2505. 07124v3 Announce Type: replace Abstract: We study inverse problems where an unknown potential is observed only through samples from the measure it induces by a convex variational principle.
By Francisco Andrade, Gabriel Peyr\'e, Clarice Poon
The paper introduces a Projected Riemannian Gradient Descent (RGD) algorithm for computing the Bures‑Wasserstein barycenter of positive definite matrices, achieving dimension‑independent linear convergence at unit step size. It resolves a previous dichotomy by showing that clipping eigenvalues to a fixed interval yields a closed‑form, non‑expansive projection in the BW metric, allowing the algorithm to match the empirical speed of unit‑step RGD while maintaining theoretical guarantees. The method also extends to the invariant matrix projection problem, providing a unified dimension‑independent analysis.
arXiv:2609. 03762v1 Announce Type: new Abstract: The computation of the Bures-Wasserstein (BW) barycenter of an ensemble of positive definite matrices arises throughout machine learning, optimal transport, and quantum information.
By A. Afham
arXiv:2607. 16384v1 Announce Type: new Abstract: For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons.
By Jingyi Zhang, Cheng Mao, Debankur Mukherjee
arXiv:2607. 02003v1 Announce Type: cross Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics.
By Matej Benko, Pierre Bousquet, Iwona Chlebicka, B{\l}a\.zej Miasojedow