arXiv:2606. 06772v1 Announce Type: cross Abstract: Understanding the generalization performance of over-parameterized neural networks has become a central topic in deep learning theory.
By Junyu Zhou, Puyu Wang, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv:2403.12187v2 Announce Type: replace-cross
Abstract: Motivated by the abundance of functional data, such as time series and images, we study the approximation and statistical learning of nonline...
By Tian-Yi Zhou, Namjoon Suh, Guang Cheng, Xiaoming Huo
arXiv:2505. 21423v3 Announce Type: replace Abstract: The remarkable generalization properties of overparameterized networks are often attributed to implicit biases, such as norm minimization at small learning rates and low sharpness in the Edge-of-Stability regime.
By Maria Matveev, Vit Fojtik, Hung-Hsu Chou, Gitta Kutyniok, Johannes Maly
arXiv:2609.31101v1 Announce Type: cross
Abstract: The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this...
By Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta
arXiv:2609.21017v1 Announce Type: cross
Abstract: We study data-driven early stopping for spectral regularisation methods in the classical non-parametric regression setting. Building on the discrepan...
By Mike Nguyen, Nicole M\"ucke
arXiv:2606. 28242v1 Announce Type: cross Abstract: Understanding how performance scales jointly with model size and data is a central problem in modern machine learning.
By Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborov\'a
arXiv:2608. 28564v1 Announce Type: cross Abstract: We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $\alpha\geq 0$ for polynomial inner-product kernels.
By Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro
arXiv:2403.04545v4 Announce Type: replace
Abstract: Scaling factors in residual branches have emerged as a prevalent method for boosting neural network performance, especially in normalization-free a...
By Zixiong Yu, Guhan Chen, Jianfa Lai, Bohan Li, Songtao Tian
arXiv:2606.25494v2 Announce Type: replace-cross
Abstract: Building on the large-sample analysis of infinitesimal gradient boosting (Dombry and Duchamps, 2024), we study the fluctuations of the proces...
By Cl\'ement Dombry, Jean-Jil Duchamps
arXiv:2604. 08625v2 Announce Type: replace-cross Abstract: We develop a theoretical framework for generalization in the interpolating regime of statistical learning.
By Gustav Olaf Yunus Laitinen-Lundstr\"om Fredriksson-Imanov
arXiv:2609.07755v1 Announce Type: new
Abstract: Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training d...
By Yuqing Wang, Ioannis G. Kevrekidis, Mikhail Belkin