arXiv Statistics ML

A functional central limit theorem for kernel gradient flow and infinitesimal gradient boosting

arXiv Machine Learning
Sep 15

Stochastic Gradient Descent over P2

The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.

By Maria Oprea, Qin Li, Yunan Yang
arXiv Machine Learning
Aug 28

Gaussian Processes and Reproducing Kernel Hilbert Spaces: Connections and Equivalences

The monograph explores the relationships between Gaussian processes and reproducing kernel Hilbert spaces (RKHS), two widely used approaches that rely on positive definite kernels. It examines how these frameworks connect and are equivalent across key topics such as regression, interpolation, numerical integration, distributional discrepancies, statistical dependence, and Gaussian process sample path properties. By establishing a unifying perspective based on the equivalence between the Gaussian Hilbert space and the RKHS, the work aims to bridge methods developed independently by the machine learning, statistics, and numerical analysis communities.

By Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, Bharath K. Sriperumbudur
arXiv Machine Learning
Sep 17

Fast Learning Rates for Physics-Informed Kernel Methods

arXiv:2609. 18901v1 Announce Type: cross Abstract: In physics-informed machine learning, a target function $u^*$ is learned from noisy value observations $y_i=u^*(x_i)+ \varepsilon_i$, together with differential information, given either by noisy observations $d_j=(Du^*)(z_j)+\xi_j$ or by a known physical constraint $Du^*=v$.

By Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti, Lorenzo Rosasco
arXiv Machine Learning
Sep 14

Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

The paper proves that stochastic gradient descent with gradient clipping and additive Gaussian noise (SGD‑CN) converges almost surely under smoothness and bounded noise assumptions, given standard decaying step sizes. The analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, showing that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods in both convex and nonconvex regimes.

By Amartya Mukherjee, Jun Liu
arXiv Machine Learning
Jul 8

A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks

arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.

By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna