arXiv AI

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

The paper introduces Normal Diffusion Dynamics Learning (NDDL), a defense framework for text-to-image diffusion models that learns normal transition dynamics from benign samples. By modeling structured, timestep‑dependent patterns across cross‑attention, latent, and noise spaces, NDDL detects backdoor attacks through deviations between observed and predicted diffusion trajectories. It also localizes triggers without prior knowledge by substituting low‑semantic words, and experiments show its effectiveness across diverse attacks.

arXiv Machine Learning
Jun 3

Backdooring Masked Diffusion Language Models

arXiv:2605. 19262v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored.

By Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou, Chengyu Huang, Pin-Yu Chen, Shengwei An
arXiv Machine Learning
Sep 14

Certifying Concept Unlearning in Text-to-Image Diffusion Models

The paper introduces a certification framework for assessing concept unlearning in text-to-image diffusion models, offering high‑confidence guarantees with bounded error on residual concept leakage. Unlike prior methods that rely solely on attack success rates from automated prompt searches, this approach combines statistical certification with worst‑case analysis along concept‑relevant embedding directions to derive explicit upper bounds on leakage probability. Evaluations across NSFW content, artistic styles, and celebrity identities reveal that certified leakage bounds exceed standard attack success rates by 16.2%, highlighting significant residual risks overlooked by existing protocols.

By Mansi, Luca Marzari, Francesco Leofante
arXiv Computer Vision
Aug 27

DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.

By Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu
arXiv Machine Learning
Aug 14

Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models

arXiv:2604. 05809v2 Announce Type: replace-cross Abstract: This paper presents Text-Guided Backdoor (TGB), an adjustable backdoor attack against multimodal pretrained models that uses natural-word triggers, namely words that can naturally occur in ordinary textual inputs.

By Yiyang Zhang, Chaojian Yu, Ziming Hong, Yuanjie Shao, Qinmu Peng, Tongliang Liu, Xinge You
arXiv Computation and Language
Sep 1

PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models

PEPPER (PErcePtion-Guided PERturbation) is a backdoor defense for text-to-image diffusion models that rewrites input captions into semantically distant yet visually similar versions, adding unobtrusive elements to disrupt trigger embeddings. By escaping the attacked neighborhood in the text embedding space, PEPPER reduces attack success, especially against text encoder-based attacks, while preserving generation quality. The method requires no training or model weight access and can be combined with existing defenses to achieve stronger, generalizable robustness.

By Oscar Chew, Po-Yi Lu, Jayden Lin, Kuan-Hao Huang, Hsuan-Tien Lin
arXiv AI
2d ago

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.

By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao