arXiv Computer Vision
Aug 31

Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models

The paper presents a comprehensive toolkit for evaluating and analyzing NSFW concept erasure in text-to-image diffusion models. It conducts the first systematic study of existing erasure methods, exploring how their mechanisms interact with empirical results. The authors offer detailed insights and practical guidance to improve content safety in diffusion models across real-world scenarios.

By Die Chen, Zhiwen Li, Cen Chen, Yuexiang Xie, Xiaodan Li, Jinyan Ye, Yingda Chen, Yaliang Li
Hugging Face Trending Papers
Jul 6

Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal.

arXiv Machine Learning
Sep 14

Certifying Concept Unlearning in Text-to-Image Diffusion Models

The paper introduces a certification framework for assessing concept unlearning in text-to-image diffusion models, offering high‑confidence guarantees with bounded error on residual concept leakage. Unlike prior methods that rely solely on attack success rates from automated prompt searches, this approach combines statistical certification with worst‑case analysis along concept‑relevant embedding directions to derive explicit upper bounds on leakage probability. Evaluations across NSFW content, artistic styles, and celebrity identities reveal that certified leakage bounds exceed standard attack success rates by 16.2%, highlighting significant residual risks overlooked by existing protocols.

By Mansi, Luca Marzari, Francesco Leofante