arXiv:2604. 09967v2 Announce Type: replace-cross Abstract: Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization.
By Ziyue Liu, Ruijie Zhang, Zhengyang Wang, Yequan Zhao, Yupeng Su, Zi Yang, Zheng Zhang
arXiv:2606. 27153v1 Announce Type: cross Abstract: Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads.
By Vincent Chen, Starrick Liu, Regis Cheng, Dance Yang, Shalfun Li, Ryan Yu, Lucy Liang, Hang Su, Roy Gan, Hao Wang, Qian Wang
arXiv:2512. 04632v2 Announce Type: replace Abstract: Orthogonality-based optimizers, such as Muon, have recently shown strong performance across large-scale training and community-driven efficiency challenges.
By Thibaut Boissin (IRIT-MISFIT), Thomas Massena (DTIPG - SNCF, IRIT-MISFIT), Franck Mamalet (IRIT-MISFIT), Mathieu Serrurier (IRIT-MISFIT)
arXiv:2608.20818v1 Announce Type: cross
Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-en...
By Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen
arXiv:2606. 16371v1 Announce Type: new Abstract: Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings.
By Bishnu Dev (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Sushil Bohara (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Martin Tak\'a\v{c} (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Samuel Horv\'ath (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE)
FedSubMuon introduces a communication‑efficient federated fine‑tuning approach for large language models by optimizing compact coefficient matrices within shared structured subspaces, thereby keeping Muon’s matrix‑aware optimization while reducing client upload size. An accuracy‑oriented variant, FedSubMuon‑GT, further adapts subspace bases using projected gradients to better align with task‑relevant directions. Experiments on instruction tuning and mathematical reasoning demonstrate that FedSubMuon‑GT achieves the best overall accuracy on most dataset‑model pairs, while FedSubMuon outperforms all matched‑budget baselines and reduces communication by up to 5.5× on Llama‑1B and 1.4× on Qwen‑4B compared to the closest baseline.
By Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang
The paper proposes a new matrix multiplication approach called SFC-CA GEMM that uses space‑filling curves to partition work in a platform‑ and shape‑oblivious way, achieving communication‑avoiding properties. It demonstrates provable asymptotic communication optimality for both square and rectangular matrices and outperforms vendor libraries on multiple x86 and Arm platforms, with speedups up to 5.5× for specific shapes and 1.8× in weighted harmonic mean throughput. The method is applied to real‑world tasks, improving large‑language‑model inference by up to 1.85× and distributed‑memory GEMM by up to 2.3× over state‑of‑the‑art frameworks.
By Evangelos Georganas, Alexander Heinecke, Pradeep Dubey
arXiv:2607. 20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale.
By Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort
arXiv:2608. 03941v1 Announce Type: new Abstract: Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm.
By Arslan Battalov, Karim Kramin, Alexander Markotenko, Sofia Sinitsina
The paper introduces SPECTRA, a runtime‑reconfigurable tiled architecture designed to accelerate speculative decoding for large language models on edge devices. SPECTRA adapts its compute engine within each tile between systolic GEMM execution and vector‑lane GEMV execution, while dynamically adjusting tile count, kernel partitioning, and communication patterns across tiles. Experiments on a 20‑tile FPGA prototype demonstrate up to a 2.09× speedup from tile‑level reconfiguration and an additional 1.25× improvement from system‑level adaptability compared to fixed designs.
By Gabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni
arXiv:2609.09676v1 Announce Type: cross
Abstract: Muon replaces matrix momentum with an approximately orthogonal polar direction, but its geometry depends on the matrix representation. For convolutio...
By Jiaxin Qing, Lexin Li
arXiv:2608. 14492v1 Announce Type: new Abstract: The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks.
By Ben Anson, Conor Houghton, Edward Milsom