arXiv:2607. 08954v1 Announce Type: cross Abstract: We study nonasymptotic convergence of primal-dual methods for a class of nonconvex constrained optimization problems with a convex-composite structure.
By Linglingzhi Zhu, Jiajin Li
arXiv:2409. 19279v2 Announce Type: replace-cross Abstract: Continuous-time models can reveal accelerated structures in distributed optimization, but their rates need not survive direct discretization.
By Kushal Chakrabarti, Mayank Baranwal
arXiv:2609.14953v1 Announce Type: cross
Abstract: This paper aims to develop new and efficient distributed algorithms for solving a class of monotone inclusions, $0 \in \sum_{i=1}^n (G_ix + T_ix)$, o...
By Nghia Nguyen-Trung, Ion Necoara, Quoc Tran-Dinh
arXiv:2507. 19895v4 Announce Type: replace-cross Abstract: In this paper, we study the distributed linear quadratic problem with fixed communication topology (DFT-LQ) and the sparse feedback linear quadratic (SF-LQ) problem through a unified optimization framework.
By Lechen Feng, Xun Li, Yuan-Hua Ni
The paper introduces Muon, an optimizer that uses a finite number of Newton‑Schulz iterations to approximate the polar factor for matrix‑valued parameters in large language model pretraining. It demonstrates that this finite iteration smooths the discontinuous polar map into a Lipschitz function of singular values, enabling a conversion from online learning regret to a stationarity guarantee in nonsmooth nonconvex optimization. The authors prove that a logarithmic depth in Newton‑Schulz suffices for convergence to stationary points, matching best‑known sample complexity bounds and extending the result to other spectral maps with similar smoothing properties.
By Mingyi Li, Taira Tsuchiya
arXiv:2502. 00470v3 Announce Type: replace-cross Abstract: Distributed empirical risk minimization (ERM) is often studied through two influential yet seemingly separate families of methods: CoCoA-type algorithms, derived from distributed dual coordinate ascent, and ADMM-type algorithms, derived from consensus and proximal splitting.
By Runxiong Wu, Andi Wang
arXiv:2609.13925v1 Announce Type: cross
Abstract: This work studies the stability and convergence of augmented primal-dual dynamics when constraint values are estimated from samples. Unbiased constra...
By Kang Liu, Mengxiao Chen, Siqi Xiong, Yi Xia
arXiv:2609.39301v1 Announce Type: cross
Abstract: This paper considers a class of multiblock nonconvex and nonsmooth optimization problems arising in many applications. Existing methods construct pro...
By Weifeng Yang
arXiv:2606. 00542v1 Announce Type: new Abstract: Shampoo-style optimizers approximate gradient covariance matrices using Kronecker-factored structures.
By Bing Liu, Wenjie Zhou, Chengcheng Zhao
arXiv:2609.00471v1 Announce Type: cross
Abstract: This paper introduces a structural taxonomy for constrained non-convex optimization based on the signature of Lagrange multipliers at KKT stationary...
By Seyed Mohsen Kazemi, Ali Movaghar, Shaahin hessabi
arXiv:2209. 15130v3 Announce Type: replace-cross Abstract: We study a general matrix optimization problem with a fixed-rank positive semidefinite (PSD) constraint.
By Yuetian Luo, Nicolas Garcia Trillos
arXiv:2607. 14731v1 Announce Type: new Abstract: Local SGD, also known as Federated Averaging, is a widely used distributed optimization algorithm.
By Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich, Aurelien Lucchi, Eduard Gorbunov, Lingxiao Wang