Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the trigger, leaving the defender at a structural disadvantage when unknown backdoors may exist in a model.
arXiv:2606. 02995v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak backdoor attacks, where adversaries poison safety alignment data to embed hidden triggers that bypass safety mechanisms.
By Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang
The paper introduces Normal Diffusion Dynamics Learning (NDDL), a defense framework for text-to-image diffusion models that learns normal transition dynamics from benign samples. By modeling structured, timestep‑dependent patterns across cross‑attention, latent, and noise spaces, NDDL detects backdoor attacks through deviations between observed and predicted diffusion trajectories. It also localizes triggers without prior knowledge by substituting low‑semantic words, and experiments show its effectiveness across diverse attacks.
By Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen
arXiv:2606. 26285v1 Announce Type: cross Abstract: Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribution target generation.
By William Aiken, Paula Branco, Guy-Vincent Jourdan, Iosif-Viorel Onut
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal repres...
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.