arXiv:2606. 07496v1 Announce Type: new Abstract: Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinator is required.
By Ming Sun, Kun Yuan
arXiv:2405. 11667v2 Announce Type: replace Abstract: Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD.
By Kumar Kshitij Patel, Margalit Glasgow, Ali Zindari, Lingxiao Wang, Sebastian U. Stich, Ziheng Cheng, Nirmit Joshi, Nathan Srebro
arXiv:2409. 19279v2 Announce Type: replace-cross Abstract: Continuous-time models can reveal accelerated structures in distributed optimization, but their rates need not survive direct discretization.
By Kushal Chakrabarti, Mayank Baranwal
arXiv:2606. 04757v1 Announce Type: cross Abstract: We study decentralized stochastic smooth convex optimization, where $M$ workers minimize an average objective using local stochastic gradients and neighbor-only communication over a fixed gossip network.
By Nitai Kluger, Amit Attia, Tomer Koren
arXiv:2603. 05774v2 Announce Type: replace Abstract: This paper addresses the distributed stochastic minimax optimization problem subject to stochastic constraints.
By Zhankun Luo, Antesh Upadhyay, Sang Bin Moon, Abolfazl Hashemi
arXiv:2607. 25492v2 Announce Type: replace Abstract: We study stochastic optimization with heavy-tailed gradient noise.
By Bin Luo, Chengchang Liu, Jonathan Allcock, Shengyu Zhang, John C. S. Lui
The paper introduces Q-MINO, a Quantization-Aware Minimal-Norm Optimizer designed to improve training of ultra-low-bit neural networks. Q-MINO uses a temporal bundle method that incorporates gradient consensus, state-drift regularization, and an alignment constraint to produce stabilized, minimum-norm update directions. The authors solve the resulting constrained subproblem with a warm-started Frank–Wolfe procedure and provide theoretical convergence guarantees via a stochastic Lyapunov Kurdyka–Łojasiewicz framework, along with numerical experiments demonstrating its effectiveness across various quantization levels.
By Don Li
arXiv:2607. 14731v1 Announce Type: new Abstract: Local SGD, also known as Federated Averaging, is a widely used distributed optimization algorithm.
By Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich, Aurelien Lucchi, Eduard Gorbunov, Lingxiao Wang
arXiv:2608. 18147v1 Announce Type: cross Abstract: Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a given input while preserving unbiasedness.
By Ran Ben Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay Vargaftik
arXiv:2608. 06563v1 Announce Type: new Abstract: Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications.
By Grigory Malinovsky
arXiv:2607. 01755v1 Announce Type: cross Abstract: In this paper, we consider the nonsmooth nonconvex decentralized optimization problem, where inter-agent communication is compressed.
By Siyuan Zhang, Nachuan Xiao, Xin Liu
arXiv:2609. 11712v1 Announce Type: cross Abstract: In this paper, we investigate the generalization performance of distributed gradient descent algorithms in a reproducing kernel Hilbert space under a robust loss function $l_{\sigma}$.
By Jun-Yi Meng, Zheng-Chu Guo, Yuan Mao