arXiv:2607. 05748v1 Announce Type: new Abstract: The community has recently developed various training-time defenses to counter neural backdoors introduced through data poisoning.
By Qi Zhao, Christian Wressnegger
The paper demonstrates that a single poisoned data point can successfully create a backdoor in linear models and ReLU neural networks without needing detailed knowledge of the training data. It establishes provable conditions under which this one‑poison attack works with high probability, achieving zero backdooring error while leaving the model’s normal performance largely unaffected. The attack relies only on coarse geometric bounds of the input space and training parameters.
By Thorsten Peinemann, Paula Arnold, Sebastian Berndt, Thomas Eisenbarth, Esfandiar Mohammadi
arXiv:2509.06896v3 Announce Type: replace
Abstract: Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training. Yet existing evalu...
By William Xu, Chenyu Zhang, Yihan Wang, Matthew Y. R. Yang, Zuoqiu Liu, Yaoliang Yu, Gautam Kamath, Yiwei Lu
arXiv:2608. 00732v1 Announce Type: new Abstract: Backdoor attacks pose a serious threat to deep neural networks, especially when training relies on third-party data, allowing adversaries to inject malicious behaviors through data poisoning.
By Zixuan Zhu, Rui Wang, Lihua Jing, Jinwen Zhong
arXiv:2507. 05113v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) are susceptible to backdoor attacks, where adversaries poison training data to implant backdoor into the victim model.
By Binyan Xu, Fan Yang, Xilin Dai, Di Tang, Kehuan Zhang
arXiv:2609.15029v1 Announce Type: cross
Abstract: Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns t...
By Aashiq Muhamed, Mona T. Diab, Virginia Smith, Andrew Ilyas, Matthew Jagielski
The paper introduces NEEDLE, a training‑free technique for removing backdoors from large language models. After a trigger is identified, NEEDLE estimates a backdoor direction and a refusal subspace using activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal‑related representations. The method requires no clean reference model or original poisoned data and achieves the lowest attack success rate and minimal impact on model performance across multiple model families and attack types.
By Minoo Kim, Vasileios Lampos, George Drayson
The paper evaluates two training‑time data poisoning attacks—label flipping and backdoor poisoning—on MNIST and Fashion‑MNIST using Logistic Regression, Linear SVM, and Random Forest classifiers. Label flipping degrades performance most for Logistic Regression and Linear SVM, while Random Forest remains relatively stable. Backdoor poisoning achieves near‑perfect attack success rates across all models while largely preserving clean‑test accuracy, highlighting the stealthy nature of targeted backdoors.
By Toshif Khan (Minot State University), Muhammad Abusaqer (Minot State University)
Backdoor attacks compromise training data so that a model retains clean accuracy but predicts an attacker-chosen target on triggered inputs. At very low poisoning rates, only a few samples convey the trigger--target association, making poison-sample selection critical.
Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the trigger, leaving the defender at a structural disadvantage when unknown backdoors may exist in a model.
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov