arXiv Machine Learning

Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels

arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.

Hugging Face Trending Papers
Aug 12

Reducing Symmetry Increase in Equivariant Neural Networks

Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries.

arXiv Machine Learning
Sep 16

Universal Feature Selection with Noisy Observations and Weak Symmetry Conditions

arXiv:2605.09396v2 Announce Type: replace-cross Abstract: This paper relaxes the restrictive symmetry conditions adopted in [4], [5] and extends their universal feature selection framework to accommo...

By Dier Tang (Department of Mathematics, The University of Hong Kong, Hong Kong, China), Guangyue Han (Department of Mathematics, The University of Hong Kong, Hong Kong, China)
arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu