The paper proves quantitative convergence to the target distribution and uniform‑in‑time propagation of chaos for Langevin‑regularized Stein variational gradient descent (SVGD). It shows that both the Stein interaction and the Langevin drift dissipate the same relative entropy, yielding exponential convergence under a log‑Sobolev inequality and providing finite‑particle entropy identities for empirical measures. Two finite‑time approaches—synchronous coupling and moving‑product entropy—are developed to give explicit Wasserstein, kernel Stein discrepancy, and total variation bounds, leading to polynomial uniform‑in‑time propagation of chaos rates.
By Sayan Banerjee, Dohyeon Kim
arXiv:2607. 14527v1 Announce Type: cross Abstract: Stein variational gradient descent (SVGD) transports interacting particles toward a target distribution through deterministic kernelized dynamics.
By Trevor Teolis, Maarten V. de Hoop
arXiv:2602. 13906v2 Announce Type: replace-cross Abstract: Stochastic approximation (SA) is a method for finding the root of an operator perturbed by noise.
By Shaan Ul Haque, Zedong Wang, Zixuan Zhang, Siva Theja Maguluri
arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.
By Jose Cribeiro-Ramallo, Florian Kalinke, Zolt\'an Szab\'o
arXiv:2609.09480v1 Announce Type: cross
Abstract: We develop Gaussian approximation bounds in higher-order Wasserstein distance $W_p$, $p\geq2$, for sums of multivariate martingale differences genera...
By Yixuan Zhang, Qiaomin Xie
arXiv:2608.29265v1 Announce Type: cross
Abstract: Kernel density estimation (KDE) is one of the most fundamental statistical estimators of density functions. Its direct implementation on a dataset of...
By Xie Wang, Nicolas Langren\'e, Wen Chen
arXiv:2602.13960v2 Announce Type: replace
Abstract: Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is ty...
By Zedong Wang, Yuyang Wang, Ijay Narang, Felix Wang, Yuzhou Wang, Siva Theja Maguluri
arXiv:2609.40193v1 Announce Type: new
Abstract: We establish near-linear accuracy bounds for the classical Moreau--Yosida unadjusted Langevin algorithm (MYULA). The target is $\pi\propto e^{-f-g}$, w...
By Yuchen Xin, Zhihua Zhang
arXiv:2607. 16384v1 Announce Type: new Abstract: For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons.
By Jingyi Zhang, Cheng Mao, Debankur Mukherjee
arXiv:2609. 26647v1 Announce Type: cross Abstract: We study statistical rates in entropic optimal transport in the semi-discrete regime where one measure has finite support and the other is subGaussian.
By Tomas Gonzalez, Gonzalo Mena
arXiv:2606. 07325v1 Announce Type: cross Abstract: We study the minimax rate of estimating a future value $\mu_{t_n+h}$ of a curve $t\mapsto\mu_t$ in the $2$-Wasserstein space $\mathcal{P}_2(\mathbb{R}^d)$ from finitely many noisy snapshots of its past, under an adiabatic bound $\|\nabla_t^k v\|\le\varepsilon$ on the $k$-th covariant derivative of the velocity field.
By Munsik Kim
arXiv:2606. 11255v1 Announce Type: new Abstract: Bernstein--Schur kernels are products of a finite-feature kernel (one with an explicit finite-dimensional feature map) and a completely monotone shift-invariant kernel: nonstationary kernels that fall between the shift-invariant and dot-product templates random features usually exploit, so in general neither Bochner sampling nor polynomial sketching applies to the full kernel directly.
By Taha Bouhsine