arXiv:2510. 14074v2 Announce Type: replace-cross Abstract: We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes.
By Elizabeth Collins-Woodfin, Inbar Seroussi
arXiv:2608. 07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally.
By Zhijun Liu, Dandan Jiang
arXiv:2606. 28573v1 Announce Type: new Abstract: Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions.
By Andrea Montanari, Kangjie Zhou
arXiv:2602.06797v3 Announce Type: replace-cross
Abstract: We study optimal learning rate (LR) schedules under the functional scaling law (FSL) framework (Li et al., 2025), which decomposes training d...
By Binghui Li, Zilin Wang, Fengling Chen, Shiyang Zhao, Ruiheng Zheng, Lei Wu
arXiv:2606. 04031v1 Announce Type: new Abstract: Coupled gradient descent--where the update of one parameter block depends on another--underlies bilevel optimization, two-time-scale stochastic approximation, and adversarial training.
By Ahanaf Hasan Ariq
arXiv:2608. 01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about.
By Gireeja Ranade, Anant Sahai