arXiv Machine Learning

RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent

arXiv Machine Learning
Sep 14

High-Probability Convergence of SGD via Batched Updates

The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.

By Feng Zhu, Robert W. Heath Jr., Aritra Mitra
arXiv Machine Learning
1d ago

On the Escaping Efficiency of Distributed Adversarial Training Algorithms

The paper compares distributed adversarial training algorithms—both centralized and decentralized—within multi‑agent learning environments. It introduces a theoretical framework to analyze how efficiently these algorithms escape local minima, a property linked to model flatness and robustness. The study finds that with small perturbation bounds and large batch sizes, decentralized methods (consensus and diffusion) escape local minima faster than centralized ones, but this advantage may diminish as attack strength increases.

By Ying Cao, Kun Yuan, Ali H. Sayed
arXiv Machine Learning
Sep 10

Approaching the Harm of Gradient Attacks While Only Flipping Labels

The paper investigates the impact of label‑flipping attacks on distributed machine learning, where an adversary can only flip a limited number of training labels. It formalizes the attack as a per‑round constrained optimization problem, derives a greedy label‑selection rule for logistic regression, and shows that this rule is provably optimal under mean aggregation. Experiments demonstrate that optimized label flipping can significantly degrade model accuracy, outperforming random flips, and that the attack transfers to other robust aggregators such as coordinate‑wise median and trimmed mean.

By Abdessamad El-Kabid, El-Mahdi El-Mhamdi
arXiv Machine Learning
Sep 23

Communication-Efficient Byzantine-Robust Federated Conformal Prediction via Partial Sharing

PRISM‑FCP is a federated conformal prediction framework that achieves Byzantine robustness while reducing communication costs. It does so by partially sharing model updates—transmitting only a subset of parameters per round—to dampen the influence of poisoned clients during training, and by filtering out suspected Byzantine clients during calibration using histogram‑based techniques. Experiments on synthetic data and UCI datasets show that PRISM‑FCP maintains near‑nominal coverage and offers favorable trade‑offs between communication overhead and predictive performance.

By Ehsan Lari, Reza Arablouei, Stefan Werner
arXiv AI
Sep 24

When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense

The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.

By Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy