arXiv Machine Learning

Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization

arXiv:2606. 07496v1 Announce Type: new Abstract: Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinator is required.

arXiv Machine Learning
Jun 3

Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness

arXiv:2509. 08726v3 Announce Type: replace-cross Abstract: This paper focuses on the decentralized stochastic optimization problem $f(\mathbf{x})=\frac{1}{m}\sum_{i=1}^m f_i(\mathbf{x})$ over a connected network of $n$ agents, where each local function has the form of $f_i(\mathbf{x}) = {\mathbb E}\left[F(\mathbf{x};{\boldsymbol \xi}_i)\right]$ which satisfies the $(L_0,L_1)$-smooth condition but possibly nonconvex and each random variable ${\boldsymbol \xi}_i$ follows distribution ${\mathcal D}_i$.

By Luo Luo, Xue Cui, Tingkai Jia, Cheng Chen
arXiv Machine Learning
Sep 17

Revisiting Distributed Sign-Based Variance Reduction

The paper addresses bias introduced by aggregating local signs in distributed sign-based variance reduction methods, which hampers optimal convergence rates. By proposing an unbiased compression of recursive gradient increments to track the global gradient at the server, the authors achieve optimal convergence rates for both nonconvex stochastic and finite-sum optimization. They provide specific rate bounds for α-norms and demonstrate matching sample complexities to centralized settings for finite-sum problems.

By Wei Jiang, Zechao Li, Lijun Zhang
arXiv Machine Learning
Sep 14

High-Probability Convergence of SGD via Batched Updates

The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.

By Feng Zhu, Robert W. Heath Jr., Aritra Mitra