arXiv Computer Vision

Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion Models

arXiv Computer Vision
Aug 31

Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models

The paper presents a comprehensive toolkit for evaluating and analyzing NSFW concept erasure in text-to-image diffusion models. It conducts the first systematic study of existing erasure methods, exploring how their mechanisms interact with empirical results. The authors offer detailed insights and practical guidance to improve content safety in diffusion models across real-world scenarios.

By Die Chen, Zhiwen Li, Cen Chen, Yuexiang Xie, Xiaodan Li, Jinyan Ye, Yingda Chen, Yaliang Li
Hugging Face Trending Papers
Jul 6

Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal.

arXiv Machine Learning
Sep 14

Certifying Concept Unlearning in Text-to-Image Diffusion Models

The paper introduces a certification framework for assessing concept unlearning in text-to-image diffusion models, offering high‑confidence guarantees with bounded error on residual concept leakage. Unlike prior methods that rely solely on attack success rates from automated prompt searches, this approach combines statistical certification with worst‑case analysis along concept‑relevant embedding directions to derive explicit upper bounds on leakage probability. Evaluations across NSFW content, artistic styles, and celebrity identities reveal that certified leakage bounds exceed standard attack success rates by 16.2%, highlighting significant residual risks overlooked by existing protocols.

By Mansi, Luca Marzari, Francesco Leofante
Hugging Face Trending Papers
Aug 18

TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion

TINA+ is a diffusion-consistent, text‑free inversion attack that probes residual visual knowledge in diffusion models after concept erasure. By using optimization‑based inversion and diffusion‑consistent trajectory regularization, it suppresses spurious trajectories that could falsely indicate retained concepts. Experiments across multiple erasure methods, tasks, and model architectures show that TINA+ reliably recovers erased concepts, revealing that many current techniques only sever text‑image links rather than eliminating underlying visual knowledge.

arXiv Computer Vision
2d ago

Continual Concept Erasure in Diffusion Models by Suppressing Cross-Edit Interference

The paper introduces CEASE, a training‑free method for continual concept erasure in text‑to‑image diffusion models. CEASE imposes two subspace constraints on a closed‑form solver to prevent interference across successive erasures, ensuring that new targets can be removed without undoing previously erased concepts. Experiments on erasing celebrities, artistic styles, and specific instances show that CEASE consistently balances erasure and preservation better than existing methods, which either degrade general generation or fail to fully erase targets.

By Yongliang Wu, Haori Lu, Jinqi Luo, Wei Cao, Xingyu Zhu, Yaoyao Liu
arXiv AI
Sep 15

DSS: Dynamic Semantic Steering for Robust Concept Erasure in Diffusion Models

The paper introduces Dynamic Semantic Steering (DSS), a training‑free, inference‑time defense for robust concept erasure in text‑to‑image diffusion models. DSS models local semantic neighborhoods geometrically, automatically identifies benign semantic anchors, and applies context‑aware, constrained feature correction using cross‑attention signals. Experiments show DSS achieves an average erasure rate of 91.0%, outperforming prior defenses while reducing semantic drift and preserving generation fidelity.

By Qinghui Gong, Zhengchun Zhou, Hua Meng, Yihuai Liang, Yuxuan Zhang