arXiv AI

SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer

SynGhost is a novel task‑agnostic backdoor attack that injects invisible syntactic backdoors into pre‑training corpora of language models. It uses an entropy‑based poisoning filter, contrastive learning to select optimal targets, and an awareness module to reduce interference between backdoors, thereby preserving the model’s pre‑training performance. Experiments demonstrate that SynGhost can transfer to multiple downstream tasks and withstand several defense mechanisms such as perplexity checks, fine‑pruning, and the maxEntropy filter.

arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
Hugging Face Trending Papers
Jul 29

ToxScreen: Detecting Whether an LLM Has Been Poisoned

As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.

arXiv Machine Learning
Jul 30

RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning

arXiv:2607. 26339v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence.

By Pushkal Kumar, Tucker Nielson, Tanish Kolhe, Shubham Zala, Vincent Li
arXiv AI
2d ago

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

The paper introduces Quarantined Expert Shutdown (QES), a new backdoor containment strategy for large language models. QES allows backdoor learning to occur during training but routes it into a designated, quarantined expert that can be disabled at deployment. The method achieves significant reductions in attack success rates while largely preserving model utility.

By Jianwei Li, Min-Seon Kim, Jung-Eun Kim