arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv:2501.14993v4 Announce Type: replace-cross
Abstract: The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent...
By Shuailong Zhu, Xiaohui Chen
arXiv:2605. 10775v2 Announce Type: replace-cross Abstract: A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity.
By Romain Petit, Clarice Poon, Gabriel Peyr\'e
The paper investigates gradient descent dynamics in the Edge of Stability regime, where a large learning rate causes persistent oscillations linked to improved generalization. It introduces a tractable continuous‑time mean–fluctuation model that couples the window‑averaged trajectory with its fluctuation covariance, derives this model rigorously from a sharp‑valley framework, and analyzes its stationary states and linear stability. The authors also extend the model to wide two‑layer networks, deriving a Wasserstein‑2 gradient flow for weights and fluctuations, proving well‑posedness, a mean‑field limit, and conditional convergence results, with numerical experiments illustrating the predictions and finite‑time limitations.
By Antonin Chodron de Courcel
arXiv:2608. 02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients.
By Jiechen Jackie Zhang, O. Deniz Akyildiz
arXiv:2505. 06589v2 Announce Type: replace-cross Abstract: Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-conditional laws, particle systems, weights of wide networks and attention patterns.
By Gabriel Peyr\'e
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2607. 04135v1 Announce Type: cross Abstract: The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points.
By Chan Li, Nigel Goldenfeld
arXiv:2606. 07120v1 Announce Type: new Abstract: Autoencoders (AEs) learn low-dimensional representations by mapping data into a latent space while minimizing reconstruction error.
By Santanu Das, Ramyak Bilas, Pascal Esser, Satyaki Mukherjee
arXiv:2510.13134v2 Announce Type: replace
Abstract: We study continuous-time dropout in controlled differential equations. We introduce a random-batch approximation of additive vector fields. On each...
By Antonio \'Alvarez-L\'opez, Mart\'in Hern\'andez
arXiv:2606. 10089v1 Announce Type: cross Abstract: In this work, we develop theoretical foundation for flow matching with neural-network-parameterized conditional velocity fields.
By Yihan He, Qishuo Yin, Yuan Cao, Jianqing Fan, Han Liu
arXiv:2605. 10792v2 Announce Type: replace-cross Abstract: We propose an implicit neural formulation of optimal transport that eliminates adversarial min--max optimization and multi-network architectures commonly used in existing approaches.
By Yesom Park, Eric Gelphman, Stanley Osher, Samy Wu Fung