arXiv AI

Quantifying and Optimizing Simplicity via Polynomial Representations

arXiv:2605. 29823v2 Announce Type: replace Abstract: Deep networks often exhibit a preference for "simple" solutions, and such a simplicity bias is widely believed to play a key role in generalization.

arXiv Machine Learning
Aug 5

Benign interpolation and Occam's razor

arXiv:2608. 03386v1 Announce Type: new Abstract: Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation.

By Tom F. Sterkenburg, Daniel A. Herrmann, Jan-Willem Romeijn
arXiv Machine Learning
Sep 14

Benign Loss Landscapes Can Coexist with Worst-Case Hardness

The paper demonstrates that tree tensor networks (TTNs) can encode arbitrary read‑once Boolean formulas, yielding polynomial‑size targets that are hard for gradient descent to learn in polynomial time, yet their loss landscapes are conditionally benign: every minimum‑norm local minimum is global. This shows that bad local minima are not the source of learning difficulty in TTNs; instead, high‑order degenerate saddle points caused by rank‑deficiency can impede learning. A case study on the parity function illustrates how TTNs can link landscape geometry to computational hardness.

By Zach Furman, Stephan W\"aldchen, Yangda Bei, Liam Hodgkinson
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu