The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.
By Maria Oprea, Qin Li, Yunan Yang
arXiv:2401. 04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom.
By Ori Shem-Ur, Yaron Oz
arXiv:2602.13960v2 Announce Type: replace
Abstract: Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is ty...
By Zedong Wang, Yuyang Wang, Ijay Narang, Felix Wang, Yuzhou Wang, Siva Theja Maguluri
arXiv:2401. 01599v4 Announce Type: replace Abstract: The generalization error curve of certain kernel regression method aims at determining the exact order of generalization error with various source condition, noise level and choice of the regularization parameter rather than the minimax rate.
By Yicheng Li, Weiye Gan, Zuoqiang Shi, Qian Lin
arXiv:2402.04691v5 Announce Type: replace-cross
Abstract: This study investigates the use of stochastic gradient descent (SGD) to learn operators between general Hilbert spaces. We study weak and str...
By Lei Shi, Jia-Qi Yang
The monograph explores the relationships between Gaussian processes and reproducing kernel Hilbert spaces (RKHS), two widely used approaches that rely on positive definite kernels. It examines how these frameworks connect and are equivalent across key topics such as regression, interpolation, numerical integration, distributional discrepancies, statistical dependence, and Gaussian process sample path properties. By establishing a unifying perspective based on the equivalence between the Gaussian Hilbert space and the RKHS, the work aims to bridge methods developed independently by the machine learning, statistics, and numerical analysis communities.
By Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, Bharath K. Sriperumbudur
arXiv:2607. 00207v1 Announce Type: cross Abstract: We develop a framework for analyzing the learning dynamics of $\ell_2$-adversarial training of single-index models on Gaussian mixtures in the high-dimensional limit under streaming stochastic gradient descent (SGD).
By Fabrizzio Sabelli
arXiv:2609. 18901v1 Announce Type: cross Abstract: In physics-informed machine learning, a target function $u^*$ is learned from noisy value observations $y_i=u^*(x_i)+ \varepsilon_i$, together with differential information, given either by noisy observations $d_j=(Du^*)(z_j)+\xi_j$ or by a known physical constraint $Du^*=v$.
By Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti, Lorenzo Rosasco
arXiv:2609.37787v1 Announce Type: new
Abstract: Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framew...
By Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou, Youzhi Zhang
The paper proves that stochastic gradient descent with gradient clipping and additive Gaussian noise (SGD‑CN) converges almost surely under smoothness and bounded noise assumptions, given standard decaying step sizes. The analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, showing that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods in both convex and nonconvex regimes.
By Amartya Mukherjee, Jun Liu
arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv:2609. 11712v1 Announce Type: cross Abstract: In this paper, we investigate the generalization performance of distributed gradient descent algorithms in a reproducing kernel Hilbert space under a robust loss function $l_{\sigma}$.
By Jun-Yi Meng, Zheng-Chu Guo, Yuan Mao