arXiv Machine Learning

Decentralized Gradient Descent: Bottleneck Regimes and Budget Complexity

arXiv:2607. 12172v1 Announce Type: cross Abstract: Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents.

arXiv Machine Learning
1d ago

SeedFlood: A Step Toward Scalable Decentralized Fine-Tuning of LLMs

SeedFlood is a novel decentralized fine‑tuning method for large language models that scales to billions of parameters and hundreds of clients. It leverages the seed‑reconstructible structure of zeroth‑order gradients to reduce message sizes to near‑zero, enabling efficient flooding across the network. Experiments show SeedFlood outperforms standard zeroth‑order baselines in communication efficiency and generalization, and rivals first‑order gossip methods while incurring far less communication cost.

By Jihun Kim, Dongyeop Lee, Namhoon Lee
arXiv Machine Learning
Sep 17

Revisiting Distributed Sign-Based Variance Reduction

The paper addresses bias introduced by aggregating local signs in distributed sign-based variance reduction methods, which hampers optimal convergence rates. By proposing an unbiased compression of recursive gradient increments to track the global gradient at the server, the authors achieve optimal convergence rates for both nonconvex stochastic and finite-sum optimization. They provide specific rate bounds for α-norms and demonstrate matching sample complexities to centralized settings for finite-sum problems.

By Wei Jiang, Zechao Li, Lijun Zhang
arXiv Machine Learning
Aug 28

A Unified Framework for Fair and Personalized Decentralized Learning under Communication Constraints

The paper introduces DMFL-SQ, a decentralized multi-task learning algorithm that integrates graph-based personalization, agnostic fairness, and compressed event-triggered communication. It provides convergence guarantees for non-convex objectives, achieving an ≠O(T^{-1/2}) stationarity rate despite sparse, quantized, and event-triggered communication, and offers PAC-Bayes generalization bounds for the fairness objective. Experiments on CIFAR-10 and the MUSMET EEG dataset show that DMFL-SQ reduces communication while preserving predictive performance and improving fairness across clients.

By Krishnendu S. Tharakan, Carlo Fischione