arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
The paper introduces FedMAST, a Federated Multi‑Axis Structural Tracing defense designed to detect and contain backdoor attacks in federated learning. FedMAST evaluates client updates through complementary structural, spectral, and historical evidence, applying tiered filtering and round‑level containment. In experiments across six backdoor attacks, FedMAST consistently achieves lower attack success rates while preserving high main‑task accuracy.
By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv:2606. 25858v1 Announce Type: cross Abstract: Federated learning is vulnerable to backdoor attacks in which malicious clients inject poisoned updates while preserving benign-task performance.
By Kavindu Herath, Joshua C. Zhao, Saurabh Bagchi
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
The paper introduces a low‑rank auditing method called LoRA as Oracle, which fits a small adapter to a hypothesis and analyzes the geometry, energy, and alignment of the resulting update relative to frozen weights. This approach directly measures what a model has internalized, independent of its output behavior, enabling detection of backdoors that behavioral audits miss. By identifying and erasing malicious internalizations within the same low‑rank subspace, the method consistently detects target classes across multiple datasets and architectures while preserving clean accuracy and operating at far lower parameter and memory cost than full‑model baselines.
By Marco Arazzi, Antonino Nocera
arXiv:2609.07147v1 Announce Type: new
Abstract: Federated learning, as a privacy-preserving distributed machine learning paradigm, faces significant threats from backdoor attacks. Compared to central...
By Jian Wang, Hong Shen, Wei Ke, Xue Hua Liu
arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.
By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
arXiv:2608.21500v1 Announce Type: cross
Abstract: Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inje...
By Yibo Peng, Long Lian, David Wagner, Sizhe Chen
arXiv:2609.40312v1 Announce Type: new
Abstract: Lossy compression is widely used in Federated Learning (FL) but is generally treated as an error source, while conventional poisoning defenses inspect...
By Sachi Shome, William Eiers
arXiv:2608. 09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it.
By Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu
The paper introduces STAIN-FL, a stealthy backdoor attack framework for federated video anomaly detection that uses natural surveillance conditions—such as low light, indoor settings, and crowd density—as contextual triggers. STAIN-FL manipulates anomaly labels and masks gradients to keep clean accuracy low while inducing trigger‑conditioned misclassification. Experiments on UCF‑Crime with I3D features show that sparse attacks remain undetectable, drop clean accuracy by less than 2%, yet achieve over 50% backdoor accuracy for hundreds of rounds under FedAvg and FedProx.
By Ashlinder Kaur, Purnima Murali Mohan, Zengxiang Li, Tram Truong-Huu
FedLNS is a server‑side framework that screens federated learning updates by representing each client’s contribution through changes in trainable normalization‑layer parameters, creating lightweight signatures that can be compared against a history‑aware cross‑client reference. The method requires no extra client‑to‑server communication, raw data, or labeled attack examples, and after screening, the remaining full‑model updates are aggregated with standard federated learning rules. Experiments on GPT‑style, BERT‑style, and LLaMA‑style models with 200 clients demonstrate that FedLNS achieves lower test perplexity than six baselines even when 40% of the population performs target manipulation under both IID and non‑IID data partitions.
By Kai Li, Jong-Ik Park, Carlee Joe-Wong, Wei Ni, Falko Dressler