The paper investigates Wasserstein-Fisher-Rao (WFR) gradient flows for sampling from probability distributions known only up to a normalisation constant. It demonstrates that for strongly log-concave targets satisfying certain curvature conditions, WFR flows preserve strong log-concavity—unlike pure Wasserstein flows, which only do so in the Gaussian case. Leveraging this property, the authors derive explicit non-asymptotic convergence rates for the symmetrised Kullback-Leibler divergence, showing an additive decomposition into Wasserstein and Fisher‑Rao contributions and eliminating the need for a warm start.
By Francesca Romana Crucinio, Sahani Pathiraja
arXiv:2411. 00214v2 Announce Type: replace-cross Abstract: Otto's Wasserstein gradient flow of the inclusive (forward) Kullback--Leibler (KL) divergence offers a principled framework for analyzing statistical inference algorithms, yet algorithms targeting the exclusive (reverse) KL divergence are rarely studied with such tools.
By Jia-Jie Zhu
arXiv:2608. 12111v1 Announce Type: cross Abstract: A novel advective Fisher-Rao metric is introduced for optimization tasks on paths of probability measures governed by the continuity equation.
By Benjamin Gess, Johannes M\"uller
arXiv:2610.02081v1 Announce Type: new
Abstract: There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often acco...
By Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona, Yen Ting Lin
arXiv:2607. 04738v1 Announce Type: cross Abstract: Reconstructing population dynamics is a central problem in the physical and data sciences.
By Markus Heinonen, Yair Shenfeld, Ricardo Baptista, Daniel Waxman, Dmitry Batenkov, Tim Cooijmans, Eli Bingham
Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein gradient flow (WGF): a curve of distributions driven by an energy functional.
This paper introduces a generative model that minimizes the second‑order Wasserstein loss (W₂) by solving a distribution‑dependent ordinary differential equation (ODE) whose dynamics involve the Kantorovich potential of the true data distribution and its current estimate. The authors prove that the time‑marginal laws of this ODE form a gradient flow for the W₂ loss, converging exponentially to the true data distribution, and propose an Euler scheme that recovers this gradient flow in the limit. An algorithm based on this scheme, combined with persistent training, is shown in experiments to outperform Wasserstein GANs in both low‑ and high‑dimensional settings when the level of persistent training is appropriately increased.
By Yu-Jui Huang, Zachariah Malik
The paper investigates a natural gradient method based on the Fisher information matrix of state-action distributions, which follows a Fisher‑Rao gradient flow within the state-action polytope under a linear potential. It establishes linear convergence rates for Fisher‑Rao gradient flows of linear programs, with the rate tied to the program’s geometry, and provides improved error bounds for entropic regularization. Additionally, the authors extend their analysis to perturbed flows, proving sublinear convergence for both perturbed Fisher‑Rao and natural gradient flows, thereby encompassing state‑action natural policy gradients.
By Johannes M\"uller, Semih \c{C}ayc{\i}, Guido Mont\'ufar
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
By Cl\'ement Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh
arXiv:2410. 01244v2 Announce Type: replace-cross Abstract: We introduce a novel Wasserstein-1 ($W_1$) path-space divergence for stochastic and deterministic dynamics and establish a Wasserstein Uncertainty Propagation (WUP) theorem that bounds the $W_1$ distance between terminal distributions by the proposed divergence, equivalently characterized by a weighted $L^2$ discrepancy between the underlying drifts and the $W_1$ distance between their initial measures.
By Ziyu Chen, Markos A. Katsoulakis, Benjamin J. Zhang
arXiv:2608.23916v1 Announce Type: new
Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. Th...
By Avinash Raju, Kai Zhang
arXiv:2606. 16610v1 Announce Type: cross Abstract: Diffusion Flow Matching (DFM) has recently emerged as a versatile framework for generative modeling, yet its theoretical convergence properties remain only partially understood.
By Marta Gentiloni Silveri, Giovanni Conforti, Alain Durmus