Hugging Face Trending Papers

Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks

Read the original on Hugging Face Trending Papers →

A common heuristic used to explain the generalization of first-order gradient methods on non-convex neural networks is that "flat interpolators generalize well" (Hochreiter and Schmidhuber, 1994; Keskar et al. , 2017), where flatness can be measured by the trace of the Hessian of the empirical loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu
arXiv Machine Learning
Aug 5

Benign interpolation and Occam's razor

arXiv:2608. 03386v1 Announce Type: new Abstract: Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation.

By Tom F. Sterkenburg, Daniel A. Herrmann, Jan-Willem Romeijn