arXiv:2606. 02328v1 Announce Type: new Abstract: We explore Riemannian optimization techniques for rank-factored matrix parameters, targeting contemporary deep learning applications.
By Nicholas Knight
arXiv:2606. 01216v1 Announce Type: new Abstract: The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling is challenging due to the presence of additional symmetries under coupled row/column scalings between the two factors.
By Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra
arXiv:2607. 19305v2 Announce Type: replace-cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations.
By Chen Ziheng
arXiv:2607. 19305v1 Announce Type: cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations.
By Chen Ziheng
arXiv:2608. 06218v1 Announce Type: cross Abstract: We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold.
By Mikhail Solonko, Molozhavenko Alexander, Maxim Rakhuba
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm.