arXiv Machine Learning

Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise

arXiv:2601. 21628v2 Announce Type: replace-cross Abstract: Diffusion models have achieved remarkable progress in image generation, but their increasing deployment raises serious concerns about privacy and copyright.

arXiv Computer Vision
Aug 27

DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.

By Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu
arXiv Machine Learning
Jul 16

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

arXiv:2607. 13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT).

By Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu
arXiv Machine Learning
Sep 14

Certifying Concept Unlearning in Text-to-Image Diffusion Models

The paper introduces a certification framework for assessing concept unlearning in text-to-image diffusion models, offering high‑confidence guarantees with bounded error on residual concept leakage. Unlike prior methods that rely solely on attack success rates from automated prompt searches, this approach combines statistical certification with worst‑case analysis along concept‑relevant embedding directions to derive explicit upper bounds on leakage probability. Evaluations across NSFW content, artistic styles, and celebrity identities reveal that certified leakage bounds exceed standard attack success rates by 16.2%, highlighting significant residual risks overlooked by existing protocols.

By Mansi, Luca Marzari, Francesco Leofante
arXiv AI
Aug 13

Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation

arXiv:2608. 11732v1 Announce Type: cross Abstract: Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intellectual property (IP) protection an increasingly critical concern when model leakage, copying, or unauthorized fine-tuning is disputed.

By Yuanmin Huang, Chen Chen, Geng Hong, Xiaoyu You, Hui Xue, Zhenxing Qian, Mi Zhang, Min Yang
arXiv AI
Jun 11

Diffusion-based Cumulative Adversarial Purification for Vision Language Models

arXiv:2506. 03933v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications.

By Jia Fu, Yongtao Wu, Yihang Chen, Kunyu Peng, Xiao Zhang, Volkan Cevher, Sepideh Pashami, Anders Holst