arXiv:2609.36610v1 Announce Type: new
Abstract: Agnostic federated learning (AFL) seeks a model that performs reliably across $m$ heterogeneous workers, but communication remains a bottleneck. We imp...
By Haomin Bai, Junyan Sun, Sifan Yang, Bo Xue, Lijun Zhang
arXiv:2606. 07496v1 Announce Type: new Abstract: Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinator is required.
By Ming Sun, Kun Yuan
The paper addresses bias introduced by aggregating local signs in distributed sign-based variance reduction methods, which hampers optimal convergence rates. By proposing an unbiased compression of recursive gradient increments to track the global gradient at the server, the authors achieve optimal convergence rates for both nonconvex stochastic and finite-sum optimization. They provide specific rate bounds for α-norms and demonstrate matching sample complexities to centralized settings for finite-sum problems.
By Wei Jiang, Zechao Li, Lijun Zhang
arXiv:2409. 19279v2 Announce Type: replace-cross Abstract: Continuous-time models can reveal accelerated structures in distributed optimization, but their rates need not survive direct discretization.
By Kushal Chakrabarti, Mayank Baranwal
arXiv:2504. 12742v2 Announce Type: replace Abstract: Decentralized Federated Learning (DFL) enables collaborative model training without relying on a central server.
By Yuan Zhou, Xinli Shi, Xuelong Li, Jiachen Zhong, Guanghui Wen, Jinde Cao
arXiv:2607. 18343v1 Announce Type: cross Abstract: Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor.
By Radhakrishna Achanta, Will Reed
arXiv:2607. 01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reconstruction overhead increase with model size and system scale.
By Mingkai Zheng, Junlin Chen, Haotian Xie, Zhao Zhang
arXiv:2302. 09832v4 Announce Type: replace Abstract: In distributed optimization and federated learning, slow and costly communication between parallel devices and the central server constitutes the primary bottleneck.
By Laurent Condat, Ivan Agarsk\'y, Grigory Malinovsky, Peter Richt\'arik
arXiv:2609.14953v1 Announce Type: cross
Abstract: This paper aims to develop new and efficient distributed algorithms for solving a class of monotone inclusions, $0 \in \sum_{i=1}^n (G_ix + T_ix)$, o...
By Nghia Nguyen-Trung, Ion Necoara, Quoc Tran-Dinh
arXiv:2606. 04757v1 Announce Type: cross Abstract: We study decentralized stochastic smooth convex optimization, where $M$ workers minimize an average objective using local stochastic gradients and neighbor-only communication over a fixed gossip network.
By Nitai Kluger, Amit Attia, Tomer Koren
SPADE-DFL is a communication‑efficient decentralized federated learning algorithm that uses a primal–dual method to allow the number of local function‑value updates between neighbor exchanges to increase with the computation budget while maintaining non‑private convergence rates. For smooth nonconvex objectives, it achieves a time‑averaged stationarity and consensus bound of ≠O(T−1/3) with only ≠Theta(T−2/3) communication rounds, where T is the number of local updates per client. The method also supports client‑level differential privacy by isolating data‑dependent increments, proving privacy for the full interactive transcript and quantifying the resulting optimization error, and demonstrates higher mean test accuracy than existing decentralized learning methods on four classification tasks.
By Mengli Wei, Mengkai Zhu, Jiawen Chen, Wenwu Yu, Duxin Che
In this paper, we consider the nonsmooth nonconvex decentralized optimization problem, where inter-agent communication is compressed. We propose a general framework that unifies various decentralized stochastic subgradient-type methods with unbiased compression and contractive compression with error compensation.