arXiv Machine Learning

Riesz-Kernel Stein Variational Gradient Descent: Renormalized Entropy and Long-Time Particle Limits

arXiv:2607. 14527v1 Announce Type: cross Abstract: Stein variational gradient descent (SVGD) transports interacting particles toward a target distribution through deterministic kernelized dynamics.

arXiv Machine Learning
Sep 1

Quantitative Target Convergence and Uniform-in-Time Propagation of Chaos for Langevin-Regularized SVGD

The paper proves quantitative convergence to the target distribution and uniform‑in‑time propagation of chaos for Langevin‑regularized Stein variational gradient descent (SVGD). It shows that both the Stein interaction and the Langevin drift dissipate the same relative entropy, yielding exponential convergence under a log‑Sobolev inequality and providing finite‑particle entropy identities for empirical measures. Two finite‑time approaches—synchronous coupling and moving‑product entropy—are developed to give explicit Wasserstein, kernel Stein discrepancy, and total variation bounds, leading to polynomial uniform‑in‑time propagation of chaos rates.

By Sayan Banerjee, Dohyeon Kim
arXiv Machine Learning
Aug 31

On two proofs of $d^2$ mixing of weighted Dikin walks

The paper investigates the mixing time of weighted Dikin walks used for sampling from exponential distributions on polytopes and truncated positive-semidefinite cones. It presents a general total-variation mixing bound under conditions of strong self-concordance, ν-symmetry, and mixed-trace regularity, achieving an “~O(d^2)" bound for polytopes and “~O(d^4)" for truncated PSD cones. A second result introduces a fourth-order bootstrap condition that yields stronger χ^2-divergence guarantees and an improved “~O(d^2)" mixing bound for a scaled Lee–Sidford metric.

By Yuansi Chen, Yunbum Kook
arXiv Machine Learning
Sep 17

Fast Learning Rates for Physics-Informed Kernel Methods

arXiv:2609. 18901v1 Announce Type: cross Abstract: In physics-informed machine learning, a target function $u^*$ is learned from noisy value observations $y_i=u^*(x_i)+ \varepsilon_i$, together with differential information, given either by noisy observations $d_j=(Du^*)(z_j)+\xi_j$ or by a known physical constraint $Du^*=v$.

By Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti, Lorenzo Rosasco