Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input.
The paper introduces "overpainting," a localized, context-aware image editing technique that allows users to specify precise or loose editing regions via a trimap. The method adapts a pretrained diffusion model with joint attention and low‑rank adaptation, incorporating attention‑dropout to balance noise, source, and mask inputs. An automated pipeline generates training data by pairing images from language‑based editing models, curating them, and extracting trimaps, enabling the model to perform a wide range of editing tasks.
By Sam Sartor, Iliyan Georgiev, Michael Fischer, Valentin Deschaintre, Pieter Peers
Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength introduces SoftPaint, a zero‑shot sampling method that uses soft masks to provide continuous, pixel‑level control over edits in diffusion models. The approach employs a Langevin‑iteration sampler that respects per‑pixel mask strengths, enabling smooth edits from preserving to fully re‑synthesizing content across image and video backbones. SoftPaint is gradient‑free, memory‑efficient, and works universally with existing diffusion models.
By Candi Zheng, Yuan Lan
arXiv:2510. 17136v2 Announce Type: replace Abstract: The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models.
By Enhao Gu, Haolin Hou
SR-Edit is a new image editing framework that uses iterative self‑refinement to improve fidelity. At each step it extracts precise, self‑consistent region separations from the model’s predictions and then enforces preservation in non‑edit areas with correction updates that stay aligned with the original sampling dynamics. Experiments show that SR‑Edit delivers better preservation and overall image quality than existing editing techniques.
By Andong Wang, Zehua Chen, Yuxuan Jiang, Jun Zhu
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions.