The paper introduces RC‑GRPO‑Editing, a region‑constrained Group Relative Policy Optimization framework for flow‑based image editing. It localizes exploration by decoupling initial noise perturbations to reduce background‑induced reward variance and adds an attention concentration reward to keep cross‑attention focused on the intended editing region. Experiments on CompBench demonstrate consistent gains in instruction adherence within the editing region while better preserving non‑target content.
By Zhuohan Ouyang, Zhe Qian, Wenhuo Cui, Chaoqun Wang
arXiv:2510. 08532v2 Announce Type: replace-cross Abstract: Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language.
By Rishubh Parihar, Or Patashnik, Daniil Ostashev, R. Venkatesh Babu, Daniel Cohen-Or, Kuan-Chieh Wang
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs.
arXiv:2609.01409v1 Announce Type: new
Abstract: Vision-language models (VLMs) have shown strong performance in generating scientific figures from text or images. However, producing publication-ready...
By Christian Greisinger, Zhixue Zhao, Steffen Eger
arXiv:2608.22780v1 Announce Type: new
Abstract: Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due...
By Qichao Ma, Jikang Cheng, Ling Liang, Zhaofei Yu, Tiejun Huang, Renye Yan
arXiv:2610.01670v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model...
By Yuan Huang, Zirui Song, Xiuying Chen
RefineEdit is a training‑free prompt‑to‑prompt image editing framework that uses a Generative Refinement Network to edit images by refining binary image codes. It couples edit localization with content generation, selecting editable positions based on signed probability differences between an editing branch and a source branch, and stabilizes edits with adaptive spatial freezing and finite bit locking. The method requires no additional training, external masks, or attention control, and outperforms other methods on PIE‑Bench in background‑preservation metrics and CLIP scores.
By Yulong Chen, Ziqian Zhang, Haoyu Zhang, Ao He, Senmao Li, Kai Wang
arXiv:2609.37492v1 Announce Type: new
Abstract: Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (...
By Zeyan Li, Wei Zhou, Hadi Amirpour, Minghao Zou, Panqi Yang, Jianfeng Xu
Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion.
"whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."
By Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh
arXiv:2606. 08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes.
By Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu
arXiv:2608. 02694v1 Announce Type: cross Abstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions.
By Lecheng Yan, Jianze Lin, Yichong Zhang, Ben Pan, Wenxi Li, Chenyang Lyu, Liting Zhou, Cathal Gurrin
arXiv:2607. 21318v1 Announce Type: cross Abstract: Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content.
By Jian Zhang, Zhijun Zhang