How Far Can Sharpness and Complexity Jointly Explain Generalization?
arXiv:2606. 29043v1 Announce Type: new Abstract: Sharpness and complexity are two central factors in the generalization analysis of deep neural networks.
arXiv:2605. 29823v2 Announce Type: replace Abstract: Deep networks often exhibit a preference for "simple" solutions, and such a simplicity bias is widely believed to play a key role in generalization.
arXiv:2606. 29043v1 Announce Type: new Abstract: Sharpness and complexity are two central factors in the generalization analysis of deep neural networks.
arXiv:2603. 00742v2 Announce Type: replace Abstract: While Adam has long been the ubiquitous default optimizer for deep neural networks, Muon has recently seen rapid adoption due to its superior training speed.
arXiv:2608. 03386v1 Announce Type: new Abstract: Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation.
arXiv:2607. 05546v1 Announce Type: cross Abstract: We develop a unified function space theory of deep fully connected neural networks.
arXiv:2605. 09609v2 Announce Type: replace Abstract: We provide counterexamples to the unimodal minimal filling architecture conjecture for polynomial neural networks (PNNs) with power activation functions.
The paper demonstrates that tree tensor networks (TTNs) can encode arbitrary read‑once Boolean formulas, yielding polynomial‑size targets that are hard for gradient descent to learn in polynomial time, yet their loss landscapes are conditionally benign: every minimum‑norm local minimum is global. This shows that bad local minima are not the source of learning difficulty in TTNs; instead, high‑order degenerate saddle points caused by rank‑deficiency can impede learning. A case study on the parity function illustrates how TTNs can link landscape geometry to computational hardness.
arXiv:2503. 07325v2 Announce Type: replace Abstract: Understanding and certifying the behavior of modern deep neural networks remains a fundamental challenge in reliable machine learning.
arXiv:2606. 00442v1 Announce Type: new Abstract: Many machine learning techniques rely on approximating a loss function's curvature, but this is notoriously hard to do at the scale of modern deep networks.
arXiv:2505. 21423v3 Announce Type: replace Abstract: The remarkable generalization properties of overparameterized networks are often attributed to implicit biases, such as norm minimization at small learning rates and low sharpness in the Edge-of-Stability regime.
arXiv:2608. 01357v1 Announce Type: new Abstract: Traditional approximation theory measures convergence rates in terms of the number of parameters or degrees of freedom.
The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.
arXiv:2606. 29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of deep learning.