arXiv AI

DeMuon: A Decentralized Muon for Matrix Optimization over Graphs

arXiv:2510. 01377v2 Announce Type: replace-cross Abstract: In this paper, we propose DeMuon, a method for decentralized matrix optimization over a given communication topology.

arXiv Machine Learning
Aug 19

Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning

The paper investigates two strategies for incorporating heterogeneous node weights in decentralized learning: embedding the weights into local losses to use a doubly stochastic matrix, and keeping the original losses while using a λ‑induced row‑stochastic matrix. By developing a weighted Hilbert‑space framework, the authors derive tighter convergence rates and show that the row‑stochastic matrix becomes self‑adjoint, reducing penalty terms that otherwise amplify consensus error. They provide conditions under which the row‑stochastic design converges faster, even with a smaller spectral gap, and offer topology‑design guidelines based on eigenvalue comparisons.

By Bing Liu, Boao Kong, Limin Lu, Kun Yuan, Chengcheng Zhao