Motion Concept Unlearning in Video Diffusion Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.23658v1 Announce Type: cross Abstract: Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While exist...
CleanVideo introduces a selective erasure framework for text-to-video diffusion models, addressing the challenge of removing undesired visual concepts from videos. The method uses a low-dimensional subspace intervention guided by a tri-modal gating mechanism that jointly considers spatiotemporal visual features, timestep signals, and textual semantics to decide where, when, and whether to intervene. Experiments on three video diffusion models demonstrate that CleanVideo effectively erases target concepts while preserving visual fidelity, temporal coherence, and outperforming existing baselines in both frame-level and video-level evaluations, even under concept-recovery attacks.
arXiv:2607. 14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult.
EraseSAE introduces a surgical concept erasure method for text-to-video diffusion models, using sparse autoencoders to decompose activations into interpretable, monosemantic features. The framework employs a contrastive attribution mechanism to isolate concept-specific kernels and applies timestep-resolved masks during inference to remove target concepts while preserving unrelated content. Experiments show that EraseSAE achieves precise, robust concept removal with minimal quality loss, outperforming existing methods.
EraseSAE is a framework for surgical concept erasure in text-to-video diffusion models. It uses a Partitioned Convolutional Sparse Autoencoder to decompose activations into interpretable sparse features, a contrastive attribution mechanism to isolate concept-specific kernels, and timestep‑resolved masks to confine erasure to active regions. Experiments show precise removal with minimal quality loss, outperforming existing methods.
FOMO is a training‑based selective video unlearning method that prioritizes preserving the original scene while removing targeted concepts. It localizes concept‑related representations for modification and employs a preservation mechanism that maintains non‑target scene information without auxiliary data. The approach extends to motion unlearning, enabling removal of concepts defined by temporal behavior, and achieves a strong balance between concept removal and scene preservation.