Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
This paper presents a zero‑shot framework for removing objects from real images using a frozen Stable Diffusion model, avoiding any task‑specific training. The pipeline combines SAM‑based mask construction, BLIP caption conditioning, DDIM inversion, background‑weighted masked null‑text optimization, decoder self‑attention masking, hard outside‑mask latent anchoring, and localized renoise‑denoise refinement. Experiments show effective removal of objects and context‑consistent replacement, with background‑weighted NTI especially helpful for complex backgrounds and repeated refinement reducing residual artifacts.
arXiv:2608. 20107v1 Announce Type: new Abstract: Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting.
Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength introduces SoftPaint, a zero‑shot sampling method that uses soft masks to provide continuous, pixel‑level control over edits in diffusion models. The approach employs a Langevin‑iteration sampler that respects per‑pixel mask strengths, enabling smooth edits from preserving to fully re‑synthesizing content across image and video backbones. SoftPaint is gradient‑free, memory‑efficient, and works universally with existing diffusion models.
MARS-CLIP is a zero‑shot semantic segmentation framework that builds on CLIP by adding a multi‑resolution feature extraction module and an attention refinement mechanism. The multi‑resolution module fuses fine‑grained local features with global context to mitigate low spatial resolution, while the attention refinement injects spatial and color biases from intermediate layers into the final self‑attention block to better recover object boundaries. Experiments on six public datasets show that MARS‑CLIP outperforms state‑of‑the‑art methods.
arXiv:2608.29243v1 Announce Type: new Abstract: Existing diffusion-based enhancement methods provide strong generative capability for low-light image enhancement (LLIE), yet they either rely on paire...
arXiv:2609.39157v1 Announce Type: new Abstract: Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generat...