arXiv AI

Dion3: Full-Stack Orthogonal Updates

arXiv:2608. 11612v1 Announce Type: cross Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step.

arXiv AI
Aug 24

Scaling Muon for Diffusion Transformers

arXiv:2608.20818v1 Announce Type: cross Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-en...

By Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen
arXiv Machine Learning
Jun 16

CacheMuon: Using Temporal Preconditioning To Approximate Polar Factor

arXiv:2606. 16371v1 Announce Type: new Abstract: Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings.

By Bishnu Dev (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Sushil Bohara (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Martin Tak\'a\v{c} (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Samuel Horv\'ath (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE)
arXiv AI
Sep 10

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

FedSubMuon introduces a communication‑efficient federated fine‑tuning approach for large language models by optimizing compact coefficient matrices within shared structured subspaces, thereby keeping Muon’s matrix‑aware optimization while reducing client upload size. An accuracy‑oriented variant, FedSubMuon‑GT, further adapts subspace bases using projected gradients to better align with task‑relevant directions. Experiments on instruction tuning and mathematical reasoning demonstrate that FedSubMuon‑GT achieves the best overall accuracy on most dataset‑model pairs, while FedSubMuon outperforms all matched‑budget baselines and reduces communication by up to 5.5× on Llama‑1B and 1.4× on Qwen‑4B compared to the closest baseline.

By Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang
arXiv AI
5d ago

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

The paper proposes a new matrix multiplication approach called SFC-CA GEMM that uses space‑filling curves to partition work in a platform‑ and shape‑oblivious way, achieving communication‑avoiding properties. It demonstrates provable asymptotic communication optimality for both square and rectangular matrices and outperforms vendor libraries on multiple x86 and Arm platforms, with speedups up to 5.5× for specific shapes and 1.8× in weighted harmonic mean throughput. The method is applied to real‑world tasks, improving large‑language‑model inference by up to 1.85× and distributed‑memory GEMM by up to 2.3× over state‑of‑the‑art frameworks.

By Evangelos Georganas, Alexander Heinecke, Pradeep Dubey
arXiv AI
Jul 24

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv:2607. 20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale.

By Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort
arXiv AI
Sep 23

SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture

The paper introduces SPECTRA, a runtime‑reconfigurable tiled architecture designed to accelerate speculative decoding for large language models on edge devices. SPECTRA adapts its compute engine within each tile between systolic GEMM execution and vector‑lane GEMV execution, while dynamically adjusting tile count, kernel partitioning, and communication patterns across tiles. Experiments on a 20‑tile FPGA prototype demonstrate up to a 2.09× speedup from tile‑level reconfiguration and an additional 1.25× improvement from system‑level adaptability compared to fixed designs.

By Gabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni