arXiv:2510. 14074v2 Announce Type: replace-cross Abstract: We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes.
By Elizabeth Collins-Woodfin, Inbar Seroussi
arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
By Zhijun Liu, Dandan Jiang
arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
By Andrea Montanari, Kangjie Zhou
arXiv:2602.06797v3 Announce Type: replace-cross
Abstract: We study optimal learning rate (LR) schedules under the functional scaling law (FSL) framework (Li et al., 2025), which decomposes training d...
By Binghui Li, Zilin Wang, Fengling Chen, Shiyang Zhao, Ruiheng Zheng, Lei Wu
arXiv:2606. 04031v1 Announce Type: new Abstract: Coupled gradient descent--where the update of one parameter block depends on another--underlies bilevel optimization, two-time-scale stochastic approximation, and adversarial training.
By Ahanaf Hasan Ariq
arXiv:2608. 01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about.
By Gireeja Ranade, Anant Sahai
arXiv:2608. 04382v1 Announce Type: new Abstract: Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization.
By Han Bao
arXiv:2609. 15047v1 Announce Type: new Abstract: Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as $1/\lambda$ in the weight decay.
By Chon-Fai Kam, Miloud Bessafi, Frederic Cadet
arXiv:2609.18901v1 Announce Type: cross
Abstract: In physics-informed machine learning, a target function $u^*$ is learned from noisy value observations $y_i=u^*(x_i)+ \varepsilon_i$, together with d...
By Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti, Lorenzo Rosasco
arXiv:2603. 04895v2 Announce Type: replace-cross Abstract: Overparameterized ML models, including neural networks, typically induce underdetermined training objectives with multiple global minima.
By Kuo-Wei Lai, Guanghui Wang, Molei Tao, Vidya Muthukumar
arXiv:2607. 03347v1 Announce Type: new Abstract: We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}.
By Joan Bruna
arXiv:2401. 01599v4 Announce Type: replace Abstract: The generalization error curve of certain kernel regression method aims at determining the exact order of generalization error with various source condition, noise level and choice of the regularization parameter rather than the minimax rate.
By Yicheng Li, Weiye Gan, Zuoqiang Shi, Qian Lin