arXiv:2606. 29519v1 Announce Type: new Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data.
By Lorenzo Livi
Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$.
The paper investigates the training dynamics of attention mechanisms in high-dimensional settings, focusing on attention-indexed models that encompass multi-layer and multi-head architectures. It shows that while the loss landscape can be described by a finite set of trace order parameters, the online stochastic gradient descent dynamics involve an infinite hierarchy of matrix moments that can be accurately approximated by a finite truncated system. The study further reveals that the choice of attention parameterization acts as an implicit bias: untied attention can get trapped in uninformative states, whereas tied attention induces symmetry breaking and enables weak recovery with θ(d² log d) samples, and untied attention exhibits a fast-slow dynamic leading to weak recovery when symmetry is broken.
By Yizhou Xu, Margarita Sagitova, Lenka Zdeborov\'a, Florent Krzakala
arXiv:2607. 07845v1 Announce Type: new Abstract: The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive.
By Marcel K\"uhn, Bernd Rosenow
arXiv:2609.01034v1 Announce Type: new
Abstract: The central flow of Cohen et al. (2025) is an empirically accurate continuous-time model of gradient descent at the edge of stability in deep learning,...
By Rapha\"el Berthier
The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.
By Anish Kataria
The paper introduces the "lift" technique for training input‑convex neural networks, replacing the traditional non‑negative weight constraint enforced by projected gradient descent or a softplus map. By adding a learnable slack variable and an unconstrained network that processes a permutation‑invariant batch summary, the lift couples batch‑dependent latent weights to the gradient, increasing update variance and enabling faster escape from the softplus shoulder. Experiments show that when the softplus method stalls at the shoulder, the lift achieves tighter fits and reconstructs targets roughly three times faster, while both methods agree when the shoulder is rarely reached.
By Ali Siahkoohi
The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.
By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands
arXiv:2607. 23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data.
By Anuj Apte
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic
arXiv:2606. 02596v1 Announce Type: new Abstract: The curvature exponent $\alpha$ in $h_k \propto \sigma_k^\alpha$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($\alpha \approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections).
By Anherutowa Calvo
arXiv:2607. 19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood.
By Jiachun Li