arXiv:2607. 09967v1 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps.
By Ethan Smith
arXiv:2607. 22444v1 Announce Type: new Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates.
By Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai
arXiv:2609.08381v1 Announce Type: cross
Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on the...
By Andrei Manolache, Mathias Niepert
arXiv:2606. 14187v1 Announce Type: new Abstract: Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation.
By Kaiwen Chen, Shuhai Zhang, Qiuwu Chen, Zimo Liu, Linxiao Li, Ying Sun, Yuchen Li, Yifan Zhang, Bo Han, Mingkui Tan
arXiv:2607. 19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood.
By Jiachun Li
arXiv:2607. 23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data.
By Anuj Apte
arXiv:2605. 18106v3 Announce Type: replace-cross Abstract: A striking geometric disparity has long persisted in the practice of deep learning.
By Tim Tsz-Kit Lau, Weijie Su
The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.
By Mingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong
The paper introduces the Drift Contract, a spectral update geometry for local learning that improves depth robustness and hyperparameter stability. By applying momentum orthogonalization with spectral step scaling to per‑layer updates, the authors achieve consistent performance across a wide range of widths and depths on CIFAR‑10 MLPs, outperforming local Adam and providing a per‑layer, input‑conditioned drift bound. The study also shows that the spectral geometry itself, rather than step‑size rules, drives the observed depth robustness, while a negative result indicates that the stability benefit is limited to non‑normalized layers.
By Fabien Polly
The paper derives an exact discrete‑time law that captures how learning‑rate schedules and weight decay interact in scale‑invariant neural networks, showing that a single scalar quantity governs the effective step size. It demonstrates that the balance point between contraction and expansion is intrinsically unstable, leading to recurrent dynamics when using constant learning rates with weight decay. The authors extend this analysis to various optimizers and datasets, confirming the law’s precision and showing that performance peaks sharply at the predicted boundary.
By Hasan Amin, Wei-Kai Chang, Rajiv Khanna
arXiv:2605.07815v2 Announce Type: replace
Abstract: Muon fixes the \emph{direction} of every matrix-valued update at the polar factor of its momentum, while each layer's step \emph{magnitude} is addr...
By Yuxuan Lou, Yang You
arXiv:2606. 17526v1 Announce Type: new Abstract: Efficient optimization is essential for training large language models.
By Da Chang, Ganzhao Yuan