arXiv:2606. 04757v1 Announce Type: cross Abstract: We study decentralized stochastic smooth convex optimization, where $M$ workers minimize an average objective using local stochastic gradients and neighbor-only communication over a fixed gossip network.
By Nitai Kluger, Amit Attia, Tomer Koren
arXiv:2509. 08726v3 Announce Type: replace-cross Abstract: This paper focuses on the decentralized stochastic optimization problem $f(\mathbf{x})=\frac{1}{m}\sum_{i=1}^m f_i(\mathbf{x})$ over a connected network of $n$ agents, where each local function has the form of $f_i(\mathbf{x}) = {\mathbb E}\left[F(\mathbf{x};{\boldsymbol \xi}_i)\right]$ which satisfies the $(L_0,L_1)$-smooth condition but possibly nonconvex and each random variable ${\boldsymbol \xi}_i$ follows distribution ${\mathcal D}_i$.
By Luo Luo, Xue Cui, Tingkai Jia, Cheng Chen
The paper addresses bias introduced by aggregating local signs in distributed sign-based variance reduction methods, which hampers optimal convergence rates. By proposing an unbiased compression of recursive gradient increments to track the global gradient at the server, the authors achieve optimal convergence rates for both nonconvex stochastic and finite-sum optimization. They provide specific rate bounds for α-norms and demonstrate matching sample complexities to centralized settings for finite-sum problems.
By Wei Jiang, Zechao Li, Lijun Zhang
arXiv:2409. 19279v2 Announce Type: replace-cross Abstract: Continuous-time models can reveal accelerated structures in distributed optimization, but their rates need not survive direct discretization.
By Kushal Chakrabarti, Mayank Baranwal
arXiv:2504. 12742v2 Announce Type: replace Abstract: Decentralized Federated Learning (DFL) enables collaborative model training without relying on a central server.
By Yuan Zhou, Xinli Shi, Xuelong Li, Jiachen Zhong, Guanghui Wen, Jinde Cao
Agnostic federated learning (AFL) seeks a model that performs reliably across $m$ heterogeneous workers, but communication remains a bottleneck. We improve communication efficiency by reducing the num...
arXiv:2607. 12172v1 Announce Type: cross Abstract: Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents.
By Nicol\`o Michelusi
arXiv:2505.20817v3 Announce Type: replace-cross
Abstract: Gradient clipping is widely used in language-model training to control heavy-tailed gradient noise and can improve convergence guarantees ove...
By Taha El Bakkali El Kadi, Savelii Chezhegov, Aleksandr Beznosikov, Samuel Horv\'ath, Eduard Gorbunov
arXiv:2609.36610v1 Announce Type: new
Abstract: Agnostic federated learning (AFL) seeks a model that performs reliably across $m$ heterogeneous workers, but communication remains a bottleneck. We imp...
By Haomin Bai, Junyan Sun, Sifan Yang, Bo Xue, Lijun Zhang
arXiv:2609.14953v1 Announce Type: cross
Abstract: This paper aims to develop new and efficient distributed algorithms for solving a class of monotone inclusions, $0 \in \sum_{i=1}^n (G_ix + T_ix)$, o...
By Nghia Nguyen-Trung, Ion Necoara, Quoc Tran-Dinh
arXiv:2607. 14731v1 Announce Type: new Abstract: Local SGD, also known as Federated Averaging, is a widely used distributed optimization algorithm.
By Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich, Aurelien Lucchi, Eduard Gorbunov, Lingxiao Wang
The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.
By Feng Zhu, Robert W. Heath Jr., Aritra Mitra