Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement
Read the original on arXiv Computer Vision →This paper presents a zero‑shot framework for removing objects from real images using a frozen Stable Diffusion model, avoiding any task‑specific training. The pipeline combines SAM‑based mask construction, BLIP caption conditioning, DDIM inversion, background‑weighted masked null‑text optimization, decoder self‑attention masking, hard outside‑mask latent anchoring, and localized renoise‑denoise refinement. Experiments show effective removal of objects and context‑consistent replacement, with background‑weighted NTI especially helpful for complex backgrounds and repeated refinement reducing residual artifacts.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.