arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv:2605. 10775v2 Announce Type: replace-cross Abstract: A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity.
By Romain Petit, Clarice Poon, Gabriel Peyr\'e
arXiv:2608. 02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients.
By Jiechen Jackie Zhang, O. Deniz Akyildiz
arXiv:2505. 06589v2 Announce Type: replace-cross Abstract: Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-conditional laws, particle systems, weights of wide networks and attention patterns.
By Gabriel Peyr\'e
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2607. 04135v1 Announce Type: cross Abstract: The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points.
By Chan Li, Nigel Goldenfeld
arXiv:2606. 07120v1 Announce Type: new Abstract: Autoencoders (AEs) learn low-dimensional representations by mapping data into a latent space while minimizing reconstruction error.
By Santanu Das, Ramyak Bilas, Pascal Esser, Satyaki Mukherjee
arXiv:2606. 10089v1 Announce Type: cross Abstract: In this work, we develop theoretical foundation for flow matching with neural-network-parameterized conditional velocity fields.
By Yihan He, Qishuo Yin, Yuan Cao, Jianqing Fan, Han Liu
arXiv:2605. 10792v2 Announce Type: replace-cross Abstract: We propose an implicit neural formulation of optimal transport that eliminates adversarial min--max optimization and multi-network architectures commonly used in existing approaches.
By Yesom Park, Eric Gelphman, Stanley Osher, Samy Wu Fung
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
By Yushun Zhang
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi