arXiv Machine Learning

Inclusive KL Gradient Flows: Otto-Wasserstein, Fisher-Rao-Gaussian, and Local-Estimator Dynamics

arXiv:2411. 00214v2 Announce Type: replace-cross Abstract: Otto's Wasserstein gradient flow of the inclusive (forward) Kullback--Leibler (KL) divergence offers a principled framework for analyzing statistical inference algorithms, yet algorithms targeting the exclusive (reverse) KL divergence are rarely studied with such tools.

arXiv Machine Learning
Sep 17

Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows

The paper investigates Wasserstein-Fisher-Rao (WFR) gradient flows for sampling from probability distributions known only up to a normalisation constant. It demonstrates that for strongly log-concave targets satisfying certain curvature conditions, WFR flows preserve strong log-concavity—unlike pure Wasserstein flows, which only do so in the Gaussian case. Leveraging this property, the authors derive explicit non-asymptotic convergence rates for the symmetrised Kullback-Leibler divergence, showing an additive decomposition into Wasserstein and Fisher‑Rao contributions and eliminating the need for a warm start.

By Francesca Romana Crucinio, Sahani Pathiraja
arXiv Machine Learning
Aug 13

Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows

arXiv:2608. 11544v1 Announce Type: cross Abstract: We propose CVaR-penalized Generative Particle Algorithm (CVaR-GPA), a robust, tail-agnostic algorithm for fine-tuning generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's tail characteristics.

By Thejani Gamage, Hyemin Gu, Zhizhen Zhang, Ziyu Chen, Markos Katsoulakis, Luc Rey-Bellet
arXiv Machine Learning
Jul 7

A Gradient Flow Perspective on Minimum MMD Estimation

arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.

By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
arXiv Machine Learning
Sep 3

Neural Variational Cut Posteriors without Upstream Data

The paper introduces NeVI‑Cut, a modular variational inference method for cut‑Bayes that does not require access to upstream data or models. It approximates the cut‑posterior by minimizing the expected downstream conditional Kullback‑Leibler divergence, using conditional normalizing flows as the variational family. The authors provide fixed‑data convergence rates, establish uniform KL approximation results for flow classes, and demonstrate the algorithm’s speed and accuracy on several applications.

By Jiafang Song, Sandipan Pramanik, Abhirup Datta
arXiv Machine Learning
Sep 14

A Generalized Tangent Approximation based Variational Inference Framework for Strongly Super-Gaussian Likelihoods

The paper introduces a new variational inference framework that uses tangent transformations to handle strongly super‑Gaussian likelihoods across a wide range of probability models. By constructing tangent minorants of the log‑likelihood through convex duality, the method achieves conjugacy with Gaussian priors, enabling tractable inference where traditional approaches struggle. The authors provide algorithmic convergence guarantees and near‑parametric risk bounds, and demonstrate superior scalability and accuracy on both simulated and real‑world datasets compared to existing variational algorithms.

By Somjit Roy, Pritam Dey, Debdeep Pati, Bani K. Mallick
arXiv Machine Learning
Jun 30

Robustness and Structure Preservation in Flow-Based Generative Models via Wasserstein Path-Space Divergences

arXiv:2410. 01244v2 Announce Type: replace-cross Abstract: We introduce a novel Wasserstein-1 ($W_1$) path-space divergence for stochastic and deterministic dynamics and establish a Wasserstein Uncertainty Propagation (WUP) theorem that bounds the $W_1$ distance between terminal distributions by the proposed divergence, equivalently characterized by a weighted $L^2$ discrepancy between the underlying drifts and the $W_1$ distance between their initial measures.

By Ziyu Chen, Markos A. Katsoulakis, Benjamin J. Zhang
arXiv Machine Learning
Aug 5

Information-Geometric Forward Policy Training in GFlowNets

arXiv:2608. 03967v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward.

By Yordan Raykov, Rodrigo Veiga