arXiv AI By Tobias Braun, Jonas Henry Grebe, Marcus Rohrbach, Anna Rohrbach

Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

Read the original on arXiv AI →

arXiv:2504. 21072v2 Announce Type: replace-cross Abstract: The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 31

Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models

The paper presents a comprehensive toolkit for evaluating and analyzing NSFW concept erasure in text-to-image diffusion models. It conducts the first systematic study of existing erasure methods, exploring how their mechanisms interact with empirical results. The authors offer detailed insights and practical guidance to improve content safety in diffusion models across real-world scenarios.

By Die Chen, Zhiwen Li, Cen Chen, Yuexiang Xie, Xiaodan Li, Jinyan Ye, Yingda Chen, Yaliang Li
arXiv AI
Sep 15

DSS: Dynamic Semantic Steering for Robust Concept Erasure in Diffusion Models

The paper introduces Dynamic Semantic Steering (DSS), a training‑free, inference‑time defense for robust concept erasure in text‑to‑image diffusion models. DSS models local semantic neighborhoods geometrically, automatically identifies benign semantic anchors, and applies context‑aware, constrained feature correction using cross‑attention signals. Experiments show DSS achieves an average erasure rate of 91.0%, outperforming prior defenses while reducing semantic drift and preserving generation fidelity.

By Qinghui Gong, Zhengchun Zhou, Hua Meng, Yihuai Liang, Yuxuan Zhang
Hugging Face Trending Papers
Jul 6

Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal.