arXiv:2606. 17196v1 Announce Type: cross Abstract: This paper is concerned with learning principal variations of random probability measures on $\mathbb{R}^m$ under the Wasserstein geometry.
By Peng Xu, Changbo Zhu, Young-Heon Kim, Xiaohui Chen
arXiv:2411. 15067v2 Announce Type: replace-cross Abstract: We investigate proximal descent methods, inspired by the minimizing movement scheme introduced by Jordan, Kinderlehrer and Otto, for optimizing entropy-regularized functionals on the Wasserstein space.
By Razvan-Andrei Lascu, Mateusz B. Majka, David \v{S}i\v{s}ka, {\L}ukasz Szpruch
arXiv:2602. 04272v2 Announce Type: replace-cross Abstract: The Importance-Weighted Evidence Lower Bound (IW-ELBO) has emerged as an effective objective for variational inference (VI), tightening the standard ELBO and mitigating the mode-seeking behaviour.
By Peiwen Jiang, Takuo Matsubara, Minh-Ngoc Tran
arXiv:2501.14993v4 Announce Type: replace-cross
Abstract: The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent...
By Shuailong Zhu, Xiaohui Chen
The paper introduces a new variational inference framework that uses tangent transformations to handle strongly super‑Gaussian likelihoods across a wide range of probability models. By constructing tangent minorants of the log‑likelihood through convex duality, the method achieves conjugacy with Gaussian priors, enabling tractable inference where traditional approaches struggle. The authors provide algorithmic convergence guarantees and near‑parametric risk bounds, and demonstrate superior scalability and accuracy on both simulated and real‑world datasets compared to existing variational algorithms.
By Somjit Roy, Pritam Dey, Debdeep Pati, Bani K. Mallick
arXiv:2506. 04480v2 Announce Type: replace-cross Abstract: This paper focuses on Geodesic Principal Component Analysis (GPCA) on a collection of probability distributions using the Otto-Wasserstein geometry.
By Nina Vesseron, Elsa Cazelles, Alice Le Brigant, Thierry Klein
arXiv:2606. 02047v1 Announce Type: cross Abstract: We introduce Convex Distance Operator Transport (CDOT), the first convex optimal transport framework that aligns distributions across heterogeneous domains by jointly preserving feature correspondence and intrinsic geometric structure.
By Junhyoung Chung, Euijong Song, Won Hwa Kim, Gunwoong Park
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
By Cl\'ement Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh
The paper investigates Wasserstein-Fisher-Rao (WFR) gradient flows for sampling from probability distributions known only up to a normalisation constant. It demonstrates that for strongly log-concave targets satisfying certain curvature conditions, WFR flows preserve strong log-concavity—unlike pure Wasserstein flows, which only do so in the Gaussian case. Leveraging this property, the authors derive explicit non-asymptotic convergence rates for the symmetrised Kullback-Leibler divergence, showing an additive decomposition into Wasserstein and Fisher‑Rao contributions and eliminating the need for a warm start.
By Francesca Romana Crucinio, Sahani Pathiraja
arXiv:2411. 00214v2 Announce Type: replace-cross Abstract: Otto's Wasserstein gradient flow of the inclusive (forward) Kullback--Leibler (KL) divergence offers a principled framework for analyzing statistical inference algorithms, yet algorithms targeting the exclusive (reverse) KL divergence are rarely studied with such tools.
By Jia-Jie Zhu
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
arXiv:2409. 18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond.
By Iskander Azangulov, George Deligiannidis, Judith Rousseau