arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv:2501.14993v4 Announce Type: replace-cross
Abstract: The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent...
By Shuailong Zhu, Xiaohui Chen
arXiv:2605. 10775v2 Announce Type: replace-cross Abstract: A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity.
By Romain Petit, Clarice Poon, Gabriel Peyr\'e
The paper investigates gradient descent dynamics in the Edge of Stability regime, where a large learning rate causes persistent oscillations linked to improved generalization. It introduces a tractable continuous‑time mean–fluctuation model that couples the window‑averaged trajectory with its fluctuation covariance, derives this model rigorously from a sharp‑valley framework, and analyzes its stationary states and linear stability. The authors also extend the model to wide two‑layer networks, deriving a Wasserstein‑2 gradient flow for weights and fluctuations, proving well‑posedness, a mean‑field limit, and conditional convergence results, with numerical experiments illustrating the predictions and finite‑time limitations.
By Antonin Chodron de Courcel
arXiv:2608. 02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients.
By Jiechen Jackie Zhang, O. Deniz Akyildiz
arXiv:2505. 06589v2 Announce Type: replace-cross Abstract: Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-conditional laws, particle systems, weights of wide networks and attention patterns.
By Gabriel Peyr\'e