arXiv Machine Learning

Spectral phase transitions and trainability in neural network learning dynamics

arXiv:2606. 28486v1 Announce Type: cross Abstract: The emergence of low-dimensional structures in the spectra of neural network weight matrices is a common empirical feature of trained models, but the dynamical origin of this phenomenon during learning remains an open problem.

arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv AI
Aug 25

FreKoo++: Learning Continuous Spectral Dynamics for Temporal Domain Generalization

FreKoo++ is a continuous spectral-dynamical framework designed for Temporal Domain Generalization (TDG). It unifies continuous Koopman modal dynamics with adaptive spectral disentanglement, mapping source-domain parameters into a latent space and modeling their evolution as a superposition of learnable continuous modes. The method handles irregular timestamps, supports arbitrary horizon extrapolation, and introduces an adaptive soft spectral weighting mechanism that isolates persistent dynamics from transient noise, achieving state‑of‑the‑art performance on discrete and continuous TDG benchmarks.

By En Yu, Xiaoyu Yang, Wei Duan, Guangquan Zhang, Jie Lu
arXiv Machine Learning
Jun 17

Eigen-Spike Emergence and Quadratic Equivalents for Conjugate Kernels on Nonlinearly Separable Data

arXiv:2605. 29669v2 Announce Type: replace-cross Abstract: Recent work in random matrix theory (RMT) has developed the notion of deterministic equivalents: typically linear surrogate models that approximate the spectral behavior of large nonlinear random matrices, such as nonlinear feature maps in neural networks (NNs).

By Collin Cranston, Zhichao Wang, Todd Kemp, Michael W. Mahoney