arXiv Machine Learning By Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou, Chengyu Huang, Pin-Yu Chen, Shengwei An

Backdooring Masked Diffusion Language Models

Read the original on arXiv Machine Learning →

arXiv:2605. 19262v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

The paper introduces Normal Diffusion Dynamics Learning (NDDL), a defense framework for text-to-image diffusion models that learns normal transition dynamics from benign samples. By modeling structured, timestep‑dependent patterns across cross‑attention, latent, and noise spaces, NDDL detects backdoor attacks through deviations between observed and predicted diffusion trajectories. It also localizes triggers without prior knowledge by substituting low‑semantic words, and experiments show its effectiveness across diverse attacks.

By Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen
Hugging Face Trending Papers
Jul 29

ToxScreen: Detecting Whether an LLM Has Been Poisoned

As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.