RefineEdit is a training‑free prompt‑to‑prompt image editing framework that uses a Generative Refinement Network to edit images by refining binary image codes. It couples edit localization with content generation, selecting editable positions based on signed probability differences between an editing branch and a source branch, and stabilizes edits with adaptive spatial freezing and finite bit locking. The method requires no additional training, external masks, or attention control, and outperforms other methods on PIE‑Bench in background‑preservation metrics and CLIP scores.
By Yulong Chen, Ziqian Zhang, Haoyu Zhang, Ao He, Senmao Li, Kai Wang
arXiv:2607. 21318v1 Announce Type: cross Abstract: Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content.
By Jian Zhang, Zhijun Zhang
The paper introduces RC‑GRPO‑Editing, a region‑constrained Group Relative Policy Optimization framework for flow‑based image editing. It localizes exploration by decoupling initial noise perturbations to reduce background‑induced reward variance and adds an attention concentration reward to keep cross‑attention focused on the intended editing region. Experiments on CompBench demonstrate consistent gains in instruction adherence within the editing region while better preserving non‑target content.
By Zhuohan Ouyang, Zhe Qian, Wenhuo Cui, Chaoqun Wang
Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content. Existing training-free editors either localize edits from terminal predictions under source and target prompts or preserve unrelated content through spatially unselective source-feature reuse without explicit region discovery.
Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion.
"whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."
By Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh
PrismGPT is a Vision‑Language Model that generates structured, region‑aware photo‑editing plans from a single image, without relying on commercial black‑box tools. It learns to diagnose aesthetic issues globally and locally while predicting precise editing parameters, using proxy‑guided learning with operation decomposition and region‑aware aesthetic ranking to bootstrap the model. A competence‑based dynamic scheduler shifts training focus from proxy tasks to the main editing task as skills improve, and all reasoning traces for fine‑tuning are self‑synthesized by the model itself. Experiments on MIT‑Adobe FiveK and a new professionally retouched benchmark, SPIRE, show PrismGPT achieves state‑of‑the‑art results using only about 6% of the training data required by previous methods.
By Ke Zhao, Hue Nguyen, Abhijith Punnappurath, Zhongling Wang, Iqbal Mohomed, Michael S. Brown
arXiv:2510. 08532v2 Announce Type: replace-cross Abstract: Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language.
By Rishubh Parihar, Or Patashnik, Daniil Ostashev, R. Venkatesh Babu, Daniel Cohen-Or, Kuan-Chieh Wang
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input.
SR-Edit is a new image editing framework that uses iterative self‑refinement to improve fidelity. At each step it extracts precise, self‑consistent region separations from the model’s predictions and then enforces preservation in non‑edit areas with correction updates that stay aligned with the original sampling dynamics. Experiments show that SR‑Edit delivers better preservation and overall image quality than existing editing techniques.
By Andong Wang, Zehua Chen, Yuxuan Jiang, Jun Zhu
MAST (Mask‑Guided Attention Control for Training‑Free Regional‑Multi Style Transfer) is a framework that enables diffusion models to apply multiple reference styles to user‑specified regions of a content image without any training or optimization. It introduces logit‑level attention mass allocation, sharpness‑aware temperature scaling, and discrepancy‑aware detail injection to address mass allocation, selectivity, and detail loss problems in regional‑multi style transfer. Experiments with two to five styles show that MAST outperforms baselines in ArtFID, FID, and R‑FID, achieving high regional style fidelity, content preservation, and scalability.
By Dongkyung Kang, Jaeyeon Hwang, Junseo Park, Minji Kang, Yeryeong Lee, Beomseok Ko, Hanyoung Roh, Jeongmin Shin, Hyeryung Jang
Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength introduces SoftPaint, a zero‑shot sampling method that uses soft masks to provide continuous, pixel‑level control over edits in diffusion models. The approach employs a Langevin‑iteration sampler that respects per‑pixel mask strengths, enabling smooth edits from preserving to fully re‑synthesizing content across image and video backbones. SoftPaint is gradient‑free, memory‑efficient, and works universally with existing diffusion models.
By Candi Zheng, Yuan Lan
arXiv:2610.01681v1 Announce Type: new
Abstract: Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-imag...
By Lidia Troeshestova, Alexander Ustyuzhanin, Sergey Kastryulin