arXiv Machine Learning

Finite-Sample Approximation of Hessian-Guided Perturbed Wasserstein Gradient Flows

The paper studies the finite‑sample approximation of a Hessian‑guided perturbed Wasserstein gradient flow (PWGF), which adds Gaussian perturbations to Wasserstein gradient descent to escape saddle points in nonconvex problems. It analyzes when an interacting‑particle approximation remains accurate over growing time horizons, showing that accumulated negative curvature can amplify errors while subsequent positive curvature can damp them. Under regularity assumptions and a fixed perturbation schedule, the authors prove high‑probability tracking bounds for both particles and objective values, construct a population‑first coupling to handle state‑dependent jumps, and verify the theory in a variance‑plus‑cosine model and a regularized matrix‑factorization setting.

arXiv Machine Learning
Sep 17

Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows

The paper investigates Wasserstein-Fisher-Rao (WFR) gradient flows for sampling from probability distributions known only up to a normalisation constant. It demonstrates that for strongly log-concave targets satisfying certain curvature conditions, WFR flows preserve strong log-concavity—unlike pure Wasserstein flows, which only do so in the Gaussian case. Leveraging this property, the authors derive explicit non-asymptotic convergence rates for the symmetrised Kullback-Leibler divergence, showing an additive decomposition into Wasserstein and Fisher‑Rao contributions and eliminating the need for a warm start.

By Francesca Romana Crucinio, Sahani Pathiraja
arXiv Machine Learning
Sep 23

Penalized Nonreversible Langevin for Constrained Sampling

The paper introduces penalized nonreversible Langevin algorithms for sampling from a target distribution constrained to a compact convex set. It combines a squared distance penalty with skew-symmetric perturbations that preserve the penalized Gibbs distribution, and provides nonasymptotic total variation and Wasserstein bounds under various smoothness and contraction assumptions. Numerical experiments demonstrate the methods on constrained Bayesian regression, classification, neural networks, and truncated sampling, highlighting acceleration in a stochastic quadratic model.

By Pervez Ali, Weihao Dong, Xiaoyu Wang
arXiv Machine Learning
Aug 13

Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows

arXiv:2608. 11544v1 Announce Type: cross Abstract: We propose CVaR-penalized Generative Particle Algorithm (CVaR-GPA), a robust, tail-agnostic algorithm for fine-tuning generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's tail characteristics.

By Thejani Gamage, Hyemin Gu, Zhizhen Zhang, Ziyu Chen, Markos Katsoulakis, Luc Rey-Bellet
Hugging Face Trending Papers
Aug 6

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures.

arXiv Machine Learning
Sep 1

Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute

The paper investigates the relationship between discrete gradient descent (GD) and its continuous-time gradient-flow counterpart in the context of ReLU neural networks. It shows that while GD states converge over a finite horizon, the exact discrete derivatives obtained via automatic differentiation do not necessarily match the derivative of the limiting flow, due to singular curvature at activation events. The authors provide a Stieltjes representation that separates continuous regional Hessians from atomic interface curvature, revealing rank-one discrepancies at activation jumps and demonstrating that even globally strongly convex residual-ReLU losses can exhibit large sensitivity ratios on certain initialization sets.

By Xiaoyang Li, Runni Zhou
arXiv AI
Oct 2

Discrete Wasserstein Flows for One-Step Generative Modeling

The paper presents a new one‑step generative modeling framework for finite state spaces, leveraging discrete Wasserstein geometry to define a target‑relative KL gradient flow over a reversible Markov kernel. The authors implement this flow at the particle level using Markov jumps and encode the resulting transport updates into a latent‑conditioned generator, enabling one‑step inference after training. Experiments on a controlled setting confirm KL dissipation, consistency between particle dynamics and probability flow, and accurate numerical scaling, while a finite‑capacity neural generator successfully tracks the exact transport targets.

By Alessandro Micheli, Andrea Zerio, Samir Bhatt
arXiv Machine Learning
Sep 15

Stochastic Gradient Descent over P2

The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.

By Maria Oprea, Qin Li, Yunan Yang