The paper presents a comprehensive toolkit for evaluating and analyzing NSFW concept erasure in text-to-image diffusion models. It conducts the first systematic study of existing erasure methods, exploring how their mechanisms interact with empirical results. The authors offer detailed insights and practical guidance to improve content safety in diffusion models across real-world scenarios.
By Die Chen, Zhiwen Li, Cen Chen, Yuexiang Xie, Xiaodan Li, Jinyan Ye, Yingda Chen, Yaliang Li
arXiv:2504. 21072v2 Announce Type: replace-cross Abstract: The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery.
By Tobias Braun, Jonas Henry Grebe, Marcus Rohrbach, Anna Rohrbach
arXiv:2607. 23492v1 Announce Type: cross Abstract: Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts.
By Shaswati Saha, Rajasekhar Anguluri, Manas Gaur
arXiv:2610.01969v1 Announce Type: new
Abstract: Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-im...
By Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu
Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal.
The paper introduces a certification framework for assessing concept unlearning in text-to-image diffusion models, offering high‑confidence guarantees with bounded error on residual concept leakage. Unlike prior methods that rely solely on attack success rates from automated prompt searches, this approach combines statistical certification with worst‑case analysis along concept‑relevant embedding directions to derive explicit upper bounds on leakage probability. Evaluations across NSFW content, artistic styles, and celebrity identities reveal that certified leakage bounds exceed standard attack success rates by 16.2%, highlighting significant residual risks overlooked by existing protocols.
By Mansi, Luca Marzari, Francesco Leofante
TINA+ is a diffusion-consistent, text‑free inversion attack that probes residual visual knowledge in diffusion models after concept erasure. By using optimization‑based inversion and diffusion‑consistent trajectory regularization, it suppresses spurious trajectories that could falsely indicate retained concepts. Experiments across multiple erasure methods, tasks, and model architectures show that TINA+ reliably recovers erased concepts, revealing that many current techniques only sever text‑image links rather than eliminating underlying visual knowledge.
arXiv:2606. 24192v1 Announce Type: cross Abstract: Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models.
By Miso Kim, Georu Lee, Yunji Kim, Hoki Kim, Jinseong Park, Woojin Lee
arXiv:2606.24192v2 Announce Type: replace-cross
Abstract: Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove...
By Miso Kim, Georu Lee, Yunji Kim, Hoki Kim, Jinseong Park, Woojin Lee
arXiv:2608. 12806v1 Announce Type: cross Abstract: The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters.
By Qiao Li, Xiaomeng Fu, Wangjia Yu, Runze He, Baisen Wang, Jiao Dai, Jizhong Han
The paper introduces CEASE, a training‑free method for continual concept erasure in text‑to‑image diffusion models. CEASE imposes two subspace constraints on a closed‑form solver to prevent interference across successive erasures, ensuring that new targets can be removed without undoing previously erased concepts. Experiments on erasing celebrities, artistic styles, and specific instances show that CEASE consistently balances erasure and preservation better than existing methods, which either degrade general generation or fail to fully erase targets.
By Yongliang Wu, Haori Lu, Jinqi Luo, Wei Cao, Xingyu Zhu, Yaoyao Liu
The paper introduces Dynamic Semantic Steering (DSS), a training‑free, inference‑time defense for robust concept erasure in text‑to‑image diffusion models. DSS models local semantic neighborhoods geometrically, automatically identifies benign semantic anchors, and applies context‑aware, constrained feature correction using cross‑attention signals. Experiments show DSS achieves an average erasure rate of 91.0%, outperforming prior defenses while reducing semantic drift and preserving generation fidelity.
By Qinghui Gong, Zhengchun Zhou, Hua Meng, Yihuai Liang, Yuxuan Zhang