Hugging Face Trending Papers

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence.

arXiv AI
Sep 4

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

LLaDA-Image is a unified framework that couples a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision‑language module based on the LLaDA2.0‑Mini diffusion language model. The approach first builds a strong visual generative prior through image‑only pre‑training and mid‑training, then fine‑tunes with a 220M‑sample generation pipeline that includes 98 real images. The resulting model produces highly photorealistic images that accurately follow fine‑grained editing instructions, and a distilled version, LLaDA‑Image‑Turbo, enables fast inference in 2–4 sampling steps. On Qwen‑Image‑Bench, LLaDA‑Image sets new state‑of‑the‑art scores for open‑source models in both English and Chinese tracks, and the authors release weights, code, and detailed recipes to support further research.

By Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie
arXiv AI
Aug 24

Scaling Muon for Diffusion Transformers

arXiv:2608.20818v1 Announce Type: cross Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-en...

By Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen
arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger