The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.
By Rub\'en Dar\'io Guerrero
The paper proposes constraining the query and key projection matrices in Transformer attention to the Stiefel manifold and optimizing them with a Riemannian Adam optimizer. It demonstrates that this geometric constraint yields significant performance gains on a CIFAR‑10 patch benchmark, with the constrained model outperforming standard AdamW by up to +6.79 percentage points. The authors also show that weight decay has no effect on the constrained frames and that the improvement is driven by a scale‑free step size rather than the manifold projection or equivariance properties.
By Rub\'en Dar\'io Guerrero
The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.
By Anish Kataria
arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.
By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey
arXiv:2606. 15551v1 Announce Type: new Abstract: The Edge of Stability (EoS) phenomenon, where gradient descent operates with sharpness exceeding the classical convergence threshold yet the loss decreases over long timescales, is ubiquitous in modern deep learning but remains poorly understood in realistic settings.
By Eric Gan
arXiv:2609.01034v1 Announce Type: new
Abstract: The central flow of Cohen et al. (2025) is an empirically accurate continuous-time model of gradient descent at the edge of stability in deep learning,...
By Rapha\"el Berthier
arXiv:2606. 14488v1 Announce Type: cross Abstract: Recent finite-time analyses of nonlinear two-time-scale stochastic approximation show that under contractive assumptions the slow iterate $Y_k$ with stepsizes $\beta_k=\Theta(k^{-1})$ and $\alpha_k=\Theta(k^{-a})$, $a\in(1/2,1)$, generally satisfies a mean-square rate of order $k^{-a}$; decoupled $k^{-1}$ rates require strong local linearity.
By Dhruv Sarkar, Vaneet Aggarwal
arXiv:2607. 17513v1 Announce Type: cross Abstract: Expert domains are trees; the Euclidean transformer is not, diluting parent-child structure exponentially at depth.
By Kwan Soo Shin, In Seok Kang, Munho Lee
arXiv:2608. 07436v1 Announce Type: new Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head.
By Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi
arXiv:2608. 19584v1 Announce Type: new Abstract: We study landscapes for complex-parameterized networks.
By Andrew Gracyk
arXiv:2608. 13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly.
By Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang
arXiv:2607. 08380v1 Announce Type: new Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian.
By Lachlan Ewen MacDonald, Ren\'e Vidal