← Back to all news
Hugging Face Trending Papers September 9, 2026

Muon-C: Operator-Aligned Muon for Convolutional Kernels

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • diffusion

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Statistics ML
Sep 10

Muon-C: Operator-Aligned Muon for Convolutional Kernels

arXiv:2609.09676v1 Announce Type: cross Abstract: Muon replaces matrix momentum with an approximately orthogonal polar direction, but its geometry depends on the matrix representation. For convolutio...

By Jiaxin Qing, Lexin Li
diffusion
More like this →
arXiv AI
Jun 9

Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning

arXiv:2604. 09967v2 Announce Type: replace-cross Abstract: Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization.

By Ziyue Liu, Ruijie Zhang, Zhengyang Wang, Yequan Zhao, Yupeng Su, Zi Yang, Zheng Zhang
llmssafety
More like this →
arXiv AI
Jul 7

Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates

arXiv:2512. 04632v2 Announce Type: replace Abstract: Orthogonality-based optimizers, such as Muon, have recently shown strong performance across large-scale training and community-driven efficiency challenges.

By Thibaut Boissin (IRIT-MISFIT), Thomas Massena (DTIPG - SNCF, IRIT-MISFIT), Franck Mamalet (IRIT-MISFIT), Mathieu Serrurier (IRIT-MISFIT)
benchmarks
More like this →
arXiv Machine Learning
Jun 3

Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering

arXiv:2606. 03899v1 Announce Type: new Abstract: Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear.

By Xianliang Li, Zihan Zhang, Weiyang Liu, Han Bao
llmssafety
More like this →
arXiv Machine Learning
Jun 26

DMuon: Efficient Distributed Muon Training with Near-Adam Overhead

arXiv:2606. 27153v1 Announce Type: cross Abstract: Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads.

By Vincent Chen, Starrick Liu, Regis Cheng, Dance Yang, Shalfun Li, Ryan Yu, Lucy Liang, Hang Su, Roy Gan, Hao Wang, Qian Wang
llmsrobotics
More like this →
arXiv AI
4d ago

Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training

arXiv:2609.36692v1 Announce Type: cross Abstract: Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram represen...

By Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu
llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea