arXiv Machine Learning

Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation

arXiv:2510. 24616v4 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks.

Hugging Face Trending Papers
Jun 22

Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions

Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models?

arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu
arXiv Machine Learning
Sep 7

Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning

The paper introduces Neural Low-Degree Filtering (Neural LoFi), a stylized limit of gradient-based training that turns hierarchical feature learning into an explicit iterative spectral procedure. In this framework, each layer independently selects directions with maximal low-degree correlation to the label, providing a tractable surrogate for deep learning and a kernel-space interpretation. Experiments on fully connected and convolutional networks show that Neural LoFi outperforms lazy random-feature baselines, recovers meaningful structured filters, and aligns with early gradient-descent feature discovery on real datasets.

By Yatin Dandi, Matteo Vilucchio, Luca Arnaboldi, Hugo Tabanelli, Florent Krzakala