The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.
By Rub\'en Dar\'io Guerrero
arXiv:2609.36252v1 Announce Type: new
Abstract: Closed-form recourse moves a rejected user along the unit gradient $\hat g$ of the classifier score $f$ by the promised distance $d_p=|f(x)|/\|\nabla f...
By Hazar Yueksel (Google)
arXiv:2605. 29547v2 Announce Type: replace-cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators.
By Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
The paper investigates how Adam’s update rule relates to natural gradient descent (NGD) by treating Adam as a diagonal empirical Fisher approximation with additional factors such as diagonal truncation, empirical label substitution, and temporal lag. Using a scale‑invariant metric, the authors quantify Adam’s geometric deviation from true NGD across four loss landscapes—well‑conditioned and ill‑conditioned linear regression, logistic regression, and a small neural network—finding that deviation is low in well‑conditioned settings but can reach about 10³ in ill‑conditioned or non‑convex scenarios. Despite higher geometric drift correlating with slower early optimization, Adam still achieves low final loss, and the improved empirical Fisher (iEF) yields more stable trajectories than the standard empirical Fisher (EF).
By Vihaan Paka-Hegde
The paper investigates the "edge of stability" phenomenon in deep learning, where Hessian eigenvalues remain stable above a classically predicted unstable threshold. It shows that many first‑order optimizers, including gradient descent, can violate this stability bound by up to a factor of 21.1, and that this deviation depends systematically on the optimizer used. The authors propose a new stability threshold based on the directional Hessian and gradient‑alignment score, which removes optimizer‑dependent offsets and offers consistent predictions while providing diagnostic tools to understand how optimizers balance temporal and spatial budgets.
By Jaerin Lee, Kyoung Mu Lee
The paper introduces the Drift Contract, a spectral update geometry for local learning that improves depth robustness and hyperparameter stability. By applying momentum orthogonalization with spectral step scaling to per‑layer updates, the authors achieve consistent performance across a wide range of widths and depths on CIFAR‑10 MLPs, outperforming local Adam and providing a per‑layer, input‑conditioned drift bound. The study also shows that the spectral geometry itself, rather than step‑size rules, drives the observed depth robustness, while a negative result indicates that the stability benefit is limited to non‑normalized layers.
By Fabien Polly
arXiv:2607. 20512v1 Announce Type: cross Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW.
By Yufeng Wang
arXiv:2608. 05136v1 Announce Type: new Abstract: Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not.
By Devender Singh
The paper proposes constraining the query and key projection matrices in Transformer attention to the Stiefel manifold and optimizing them with a Riemannian Adam optimizer. It demonstrates that this geometric constraint yields significant performance gains on a CIFAR‑10 patch benchmark, with the constrained model outperforming standard AdamW by up to +6.79 percentage points. The authors also show that weight decay has no effect on the constrained frames and that the improvement is driven by a scale‑free step size rather than the manifold projection or equivariance properties.
By Rub\'en Dar\'io Guerrero
arXiv:2606. 25971v1 Announce Type: new Abstract: Modern neural network training relies on optimizers such as Adam and Muon which act on each weight matrix as a single object.
By Alexander H\"agele, Alejandro Hern\'andez-Cano, Atli Kosson, Martin Jaggi
arXiv:2607. 16612v1 Announce Type: cross Abstract: Backpropagation makes training deep networks memory intensive because it must store intermediate activations.
By Tian Qin, Wei-Min Huang
arXiv:2606. 29119v1 Announce Type: cross Abstract: We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle outer loop over neural-network parameters or structure is worth building.
By Ramchand Kumaresan