arXiv AI

When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense

The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.

arXiv Machine Learning
Aug 20

FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs

FedLNS is a server‑side framework that screens federated learning updates by representing each client’s contribution through changes in trainable normalization‑layer parameters, creating lightweight signatures that can be compared against a history‑aware cross‑client reference. The method requires no extra client‑to‑server communication, raw data, or labeled attack examples, and after screening, the remaining full‑model updates are aggregated with standard federated learning rules. Experiments on GPT‑style, BERT‑style, and LLaMA‑style models with 200 clients demonstrate that FedLNS achieves lower test perplexity than six baselines even when 40% of the population performs target manipulation under both IID and non‑IID data partitions.

By Kai Li, Jong-Ik Park, Carlee Joe-Wong, Wei Ni, Falko Dressler
Hugging Face Trending Papers
Jun 13

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.

arXiv Machine Learning
Sep 23

FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in Federated Learning

FedNIA is a defense framework for federated learning that identifies and excludes malicious clients without needing a central test dataset. It works by injecting random noise inputs and analyzing layerwise activation patterns with an autoencoder to detect abnormal behaviors caused by data poisoning. The method can counter various attack types—including sample poisoning, label flipping, and backdoors—even when multiple attackers collaborate, and shows strong performance on non‑iid federated datasets.

By Ehsan Hallaji, Roozbeh Razavi-Far, Mehrdad Saif
arXiv Machine Learning
Sep 7

Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates

The paper introduces FedIoC, a federated learning framework that embeds structured threat indicators into gradient updates using a supervised contrastive loss. By aligning gradients from clients that share indicators for the same attack campaign, the server can cluster updates via cosine similarity to recover global campaign patterns without transmitting sensitive indicators. Experiments on two public threat‑detection benchmarks show that the server successfully identifies cross‑organizational campaign cohorts from fragmented local data.

By Manuel R\"oder, Bibin Babu, Frank-Michael Schleif
arXiv Machine Learning
Aug 27

Rethinking the Transferable Adversarial Attacks and Robust Defense in Federated Learning

The paper investigates how adversarial examples transfer between client models in federated learning and explores the relationship between these examples and client data distributions. It proposes a defense strategy based on adversarial training that leverages the transferability of model robustness. Experiments on real-life datasets demonstrate that the new attack and defense methods outperform existing state‑of‑the‑art approaches.

By Zuobin Xiong, Deval Mukherjee, Homook Cho, Wei Li