arXiv:2606. 16899v1 Announce Type: new Abstract: Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when using standard constant decoupled weight decay.
By Kaiyue Wen, Xingyu Dang, Kaifeng Lyu, Tengyu Ma, Percy Liang
arXiv:2606. 25971v1 Announce Type: new Abstract: Modern neural network training relies on optimizers such as Adam and Muon which act on each weight matrix as a single object.
By Alexander H\"agele, Alejandro Hern\'andez-Cano, Atli Kosson, Martin Jaggi
The paper introduces a physical response-and-memory model for the Muon optimizer, explaining its semi‑orthogonalized momentum update as the maximally dissipative direction under an output‑side safety budget. It treats the weight matrix as a responsive medium with internal stress, showing that momentum corresponds to accumulated stress whose relaxation occurs over multiple timescales—fast and slow. Based on this, the authors propose the Bi‑Maxwell optimizer, which uses a two‑timescale memory kernel and achieves target loss in fewer steps on a public large‑language‑model benchmark.
By Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu
arXiv:2609.07017v1 Announce Type: new
Abstract: Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. Alth...
By Jinghui Yuan, Hongtao Zhang, Jade Zou, Tianyu Li, Wenjie Zhou, Tianyu He, Wei Chen
arXiv:2608.21024v1 Announce Type: new
Abstract: Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight,...
By Guoxiang Xu, Bince Qu, Qi Sun, Cheng Zhuo
arXiv:2609.25914v1 Announce Type: new
Abstract: Complex-valued neural networks (CVNNs) are increasingly adopted for complex-valued data; however, they are often trained with first-order optimizers in...
By Enrico Ballini, Allan Peter Engsig-Karup, Tito Andriollo