arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).
By Steffen Dereich, Arnulf Jentzen, Adrian Riekert
arXiv:2607. 00207v1 Announce Type: cross Abstract: We develop a framework for analyzing the learning dynamics of $\ell_2$-adversarial training of single-index models on Gaussian mixtures in the high-dimensional limit under streaming stochastic gradient descent (SGD).
By Fabrizzio Sabelli
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
The paper introduces BROT, a two‑step approach for estimating optimal transport maps. First, it computes the unregularized OT plan, then fits a deep neural network to the resulting barycentric targets using least‑squares regression. The authors prove that, under standard regularity conditions, BROT achieves the minimax convergence rate when the true OT map is Lipschitz, and demonstrate its effectiveness on synthetic data, images, and downstream tasks such as single‑cell perturbation prediction and unsupervised domain adaptation.
By Kunwoong Kim, Insung Kong, Yongdai Kim
arXiv:2608. 09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success.
By Binchuan Qi
arXiv:2608. 11544v1 Announce Type: cross Abstract: We propose CVaR-penalized Generative Particle Algorithm (CVaR-GPA), a robust, tail-agnostic algorithm for fine-tuning generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's tail characteristics.
By Thejani Gamage, Hyemin Gu, Zhizhen Zhang, Ziyu Chen, Markos Katsoulakis, Luc Rey-Bellet
arXiv:2410. 02596v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algorithms.
By Rui Hu, Yifan Zhang, Zhuoran Li, Longbo Huang
The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.
By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv:2606. 27171v1 Announce Type: new Abstract: This work addresses the problem of variance in stochastic gradient estimation for machine learning optimization.
By Jonne Pohjankukka, Jukka Heikkonen
arXiv:2510.17072v2 Announce Type: replace
Abstract: Regression with non-Euclidean responses---e.g., probability distributions, networks, symmetric positive-definite matrices, and compositions---has b...
By Kyum Kim, Yaqing Chen, Paromita Dubey
arXiv:2606. 06418v1 Announce Type: new Abstract: Many modern applications of deep learning involve training a neural network via a one-step prediction loss (e.
By Thomas T. Zhang, Alok Shah, Yifei Zhang, Vincent Zhang, Nikolai Matni, Max Simchowitz
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager