arXiv Machine Learning By Xiaohe Jiang (University of Exeter), Guoqiang Zhang (University of Exeter), Tianjin Huang (University of Exeter), Ronghui Mu (University of Exeter)

NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers

Read the original on arXiv Machine Learning →

The paper introduces Newton-Schulz Attention (NS-Attn.), a parameter‑free transformation that applies a finite Newton‑Schulz polynomial step to the output of each attention head in Vision Transformers. By normalizing each head’s feature‑by‑token matrix with its Frobenius norm, applying the NS step, and restoring the norm, the method aims to reduce spectral concentration and increase effective rank before merging heads. Experiments on ViT and Swin models over CIFAR‑10 and CIFAR‑100 show consistent accuracy gains of 0.25–0.83 percentage points, though with added inference latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 23

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.

By Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai
arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger
arXiv Machine Learning
Sep 22

COREM: Cosine-Relation Momentum Reshaping with Stateful Writeback

The paper introduces COREM, a Cosine-Relation Momentum Reshaping method that exploits relational structure within matrix‑valued optimizer states. COREM partitions the momentum state into update units, computes cosine relations among them, and reshapes the momentum before writing it back, thereby influencing both current and future optimization dynamics. Experiments on CIFAR‑10 and enwik8 show that COREM improves mid‑to‑late training performance and enhances spectral properties while using fewer FLOPs than the Muon baseline.

By Yan Wang, Xiaochuan Wang, Yuxiang Sun