arXiv AI By Jiahe Yan, Pratik Chaudhari, Leonard Kleinrock

dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD

Read the original on arXiv AI →

arXiv:2412. 07151v1 Announce Type: cross Abstract: Distributed model training needs to be adapted to challenges such as the straggler effect and Byzantine attacks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
2d ago

Robustifying Asynchronous SGD via Soft Throttling

arXiv:2609.39357v1 Announce Type: new Abstract: Asynchronous SGD is a popular algorithm for distributed learning where each client's gradient update is applied on arrival. This leads to a speed-up, b...

By Kaoru Otsuka, Maxime Meyer, Yuki Takezawa, Makoto Yamada, Anastasia Koloskova
arXiv Machine Learning
Sep 4

A Nesterov-Accelerated Byzantine-Robust Federated Learning

The paper proposes Byrd-NAFL, a Byzantine‑robust federated learning algorithm that incorporates Nesterov’s momentum and resilient aggregation rules. It achieves fast and safe convergence under non‑convex, smooth loss functions with relaxed gradient assumptions, and provides a finite‑time convergence guarantee. Experiments show that Byrd-NAFL outperforms existing methods in convergence speed, accuracy, and resilience to various malicious attacks.

By Lihan Xu, Xiaoyi Fan, Gang Wang, Runhao Zeng, Xiping Hu, Yanjie Dong
arXiv AI
Sep 24

When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense

The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.

By Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy