arXiv:2511. 07210v3 Announce Type: replace-cross Abstract: Clean-image backdoor attacks, which use only label manipulation in training datasets to compromise deep neural networks, pose a significant threat to security-critical applications.
By Binyan Xu, Fan Yang, Di Tang, Xilin Dai, Kehuan Zhang
The paper introduces Normal Diffusion Dynamics Learning (NDDL), a defense framework for text-to-image diffusion models that learns normal transition dynamics from benign samples. By modeling structured, timestep‑dependent patterns across cross‑attention, latent, and noise spaces, NDDL detects backdoor attacks through deviations between observed and predicted diffusion trajectories. It also localizes triggers without prior knowledge by substituting low‑semantic words, and experiments show its effectiveness across diverse attacks.
By Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal repres...
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave like backdoor-like triggers without being maliciously inserted.
arXiv:2607. 05516v1 Announce Type: cross Abstract: Model-specific adversarial attacks have been extensively studied.
By Paul K. Mandal, Pavan Reddy, Tristan Malatynski
arXiv:2607. 05516v2 Announce Type: replace-cross Abstract: Model-specific adversarial attacks have been extensively studied.
By Paul K. Mandal, Pavan Reddy, Tristan Malatynski