arXiv Machine Learning By Harsh Vardhan, Hossein Taheri, Arya Mazumdar

Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks

Read the original on arXiv Machine Learning →

arXiv:2606. 04429v1 Announce Type: cross Abstract: A common heuristic used to explain the generalization of first-order gradient methods on non-convex neural networks is that "flat interpolators generalize well" (Hochreiter and Schmidhuber, 1994; Keskar et al.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu
arXiv Statistics ML
3d ago

Grokking through the Lens of Minimum-Norm Interpolation

The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization. whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."

By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli