A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable.
arXiv:2606. 29176v1 Announce Type: new Abstract: A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation.
By Tejas Pradeep Shirodkar
The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.
By Rub\'en Dar\'io Guerrero
arXiv:2607. 10203v2 Announce Type: replace-cross Abstract: Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better predictions and can be routed adaptively.
By Achyuthan Sivasankar
arXiv:2606. 01090v1 Announce Type: cross Abstract: Equivariance theory predicts that an architectural symmetry prior reduces sample complexity by a factor of |G|; this is widely cited but rarely measured as a scaling law with controls that separate the prior from its confounds.
By Ahmed M. Adly
arXiv:2606. 13092v3 Announce Type: replace Abstract: Scale buys interpolation; structure buys certifiable transfer.
By Hongbo Wang
arXiv:2607. 10203v1 Announce Type: cross Abstract: Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better predictions and can be routed adaptively.
By Achyuthan Sivasankar
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le
The paper proposes constraining the query and key projection matrices in Transformer attention to the Stiefel manifold and optimizing them with a Riemannian Adam optimizer. It demonstrates that this geometric constraint yields significant performance gains on a CIFAR‑10 patch benchmark, with the constrained model outperforming standard AdamW by up to +6.79 percentage points. The authors also show that weight decay has no effect on the constrained frames and that the improvement is driven by a scale‑free step size rather than the manifold projection or equivariance properties.
By Rub\'en Dar\'io Guerrero
arXiv:2609.08381v1 Announce Type: cross
Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on the...
By Andrei Manolache, Mathias Niepert
arXiv:2609.01129v1 Announce Type: new
Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under compos...
By Jiming Feng, Junliang Li
The paper introduces the Latent Generative Solver (LGS), a neural PDE solver that combines a Physics VAE, a Pyramidal Flow-Forcing Transformer, and input noising to achieve generalization across twelve PDE families and stable long-term rollouts. LGS matches or surpasses deterministic baselines on one-step predictions, outperforms them on 5- and 10-step rollouts, and significantly reduces long-horizon error while cutting compute costs. It also adapts efficiently to unseen higher-resolution systems, demonstrating strong empirical performance on 2D regular-grid PDE simulations.
By Zituo Chen, Sili Deng