arXiv:2506. 04480v2 Announce Type: replace-cross Abstract: This paper focuses on Geodesic Principal Component Analysis (GPCA) on a collection of probability distributions using the Otto-Wasserstein geometry.
By Nina Vesseron, Elsa Cazelles, Alice Le Brigant, Thierry Klein
arXiv:2311. 15365v3 Announce Type: replace Abstract: We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space.
By Noboru Isobe
Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein gradient flow (WGF): a curve of distributions driven by an energy functional.
arXiv:2607. 04738v1 Announce Type: cross Abstract: Reconstructing population dynamics is a central problem in the physical and data sciences.
By Markus Heinonen, Yair Shenfeld, Ricardo Baptista, Daniel Waxman, Dmitry Batenkov, Tim Cooijmans, Eli Bingham
arXiv:2510. 04602v4 Announce Type: replace-cross Abstract: Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space.
By Eduardo Fernandes Montesuma, Yassir Bendou, Mike Gartrell
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2608. 01434v1 Announce Type: new Abstract: Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions.
By Srinivasa Rao P Vangmayi P Reddy
arXiv:2606. 10089v1 Announce Type: cross Abstract: In this work, we develop theoretical foundation for flow matching with neural-network-parameterized conditional velocity fields.
By Yihan He, Qishuo Yin, Yuan Cao, Jianqing Fan, Han Liu
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
arXiv:2501. 07400v2 Announce Type: replace-cross Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precise sense, adapted to the coordinate system distinguished by the activations.
By Thomas Chen
arXiv:2606. 27767v1 Announce Type: new Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning.
By Cl\'ement Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh
arXiv:2607. 02003v1 Announce Type: cross Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics.
By Matej Benko, Pierre Bousquet, Iwona Chlebicka, B{\l}a\.zej Miasojedow