RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 23492v1 Announce Type: cross Abstract: Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts.
Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of damaging semantically related non-target concepts. Existing value-space methods remove the component of each cross-attention value along the target concept direction, implicitly treating target identity and shared visual structure as the same signal.
arXiv:2512. 02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations arise over time.
arXiv:2607. 06432v1 Announce Type: cross Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to suppress unwanted concepts after training.
STAGE is a training‑free, closed‑form framework for concept erasure in native text‑to‑3D generators. It treats erasure as a stage‑aware editing problem, applying low‑dimensional affine corrections separately to the structural and appearance stages of the pipeline. Experiments on the TRELLIS generator show that STAGE outperforms adapted baselines, achieving a composite score of 66.7 versus 53.2 across 15 shape, material, and object concepts.
The paper introduces CEASE, a training‑free method for continual concept erasure in text‑to‑image diffusion models. CEASE imposes two subspace constraints on a closed‑form solver to prevent interference across successive erasures, ensuring that new targets can be removed without undoing previously erased concepts. Experiments on erasing celebrities, artistic styles, and specific instances show that CEASE consistently balances erasure and preservation better than existing methods, which either degrade general generation or fail to fully erase targets.