The paper introduces Quarantined Expert Shutdown (QES), a new backdoor containment strategy for large language models. QES allows backdoor learning to occur during training but routes it into a designated, quarantined expert that can be disabled at deployment. The method achieves significant reductions in attack success rates while largely preserving model utility.
By Jianwei Li, Min-Seon Kim, Jung-Eun Kim
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2608.30041v1 Announce Type: cross
Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later pri...
By Wujie Xiong, Rabimba Karanjai, Yang Lu, Weidong Shi, Lei Xu
The paper investigates how inference optimization for large language models can introduce numerical inconsistencies that trigger hidden backdoors. It introduces two types of optimization‑triggered backdoors: the Input‑Specific Optimization Backdoor (ISOB) and the Universal Optimization Backdoor (UOB), the latter enabling a model to remain benign under normal execution but activate a backdoor when optimization is applied. Experiments on seven open‑source LLMs, across multiple tasks and optimization backends, show UOB can achieve up to 100% attack success while maintaining clean accuracy, and the authors propose three defenses that reduce the attack success rate to 0.02.
By Yifei Wang, Yida Yang, Tianlin Li, Xiaohan Zhang, Xiaoyu Zhang, Li Pan
arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.
By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
arXiv:2607. 19490v1 Announce Type: cross Abstract: Peer-to-peer distributed inference executes a Large Language Model (LLM) on pooled consumer hardware by spreading its layers across many nodes.
By Mert Cihangiroglu, Antonino Nocera
arXiv:2606. 12703v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions.
By Tarun Sharma
FedLNS is a server‑side framework that screens federated learning updates by representing each client’s contribution through changes in trainable normalization‑layer parameters, creating lightweight signatures that can be compared against a history‑aware cross‑client reference. The method requires no extra client‑to‑server communication, raw data, or labeled attack examples, and after screening, the remaining full‑model updates are aggregated with standard federated learning rules. Experiments on GPT‑style, BERT‑style, and LLaMA‑style models with 200 clients demonstrate that FedLNS achieves lower test perplexity than six baselines even when 40% of the population performs target manipulation under both IID and non‑IID data partitions.
By Kai Li, Jong-Ik Park, Carlee Joe-Wong, Wei Ni, Falko Dressler
arXiv:2607. 23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly.
By Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern
arXiv:2606. 26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions.
By Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi, Jayaram Kumarapu
DecoyTrace is a proactive cyber‑deception defense designed for strictly serverless decentralized federated learning (DFL). It deploys mobile DecoyNodes that generate chaotic decoy challenges, disseminate dual models (clean vs. decoy) based on neighbor trust, and use three‑state semantic metrics to isolate malicious sources and recover models. Across sixty configurations on the NEBULA platform, DecoyTrace restores model utility with minimal performance loss and reduces CPU and network usage by up to two‑thirds.
By Pedro Beltr\'an-L\'opez, Enrique Tom\'as Mart\'inez Beltr\'an, Pantaleone Nespoli, Manuel Gil P\'erez, Alberto Huertas Celdr\'an