Attention-Scoped Guidance: Training-Free Spatial Control for Image Editing
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
RefineEdit is a training‑free prompt‑to‑prompt image editing framework that uses a Generative Refinement Network to edit images by refining binary image codes. It couples edit localization with content generation, selecting editable positions based on signed probability differences between an editing branch and a source branch, and stabilizes edits with adaptive spatial freezing and finite bit locking. The method requires no additional training, external masks, or attention control, and outperforms other methods on PIE‑Bench in background‑preservation metrics and CLIP scores.
arXiv:2607. 21318v1 Announce Type: cross Abstract: Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content.
The paper introduces RC‑GRPO‑Editing, a region‑constrained Group Relative Policy Optimization framework for flow‑based image editing. It localizes exploration by decoupling initial noise perturbations to reduce background‑induced reward variance and adds an attention concentration reward to keep cross‑attention focused on the intended editing region. Experiments on CompBench demonstrate consistent gains in instruction adherence within the editing region while better preserving non‑target content.
Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content. Existing training-free editors either localize edits from terminal predictions under source and target prompts or preserve unrelated content through spatially unselective source-feature reuse without explicit region discovery.
Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion. "whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."
PrismGPT is a Vision‑Language Model that generates structured, region‑aware photo‑editing plans from a single image, without relying on commercial black‑box tools. It learns to diagnose aesthetic issues globally and locally while predicting precise editing parameters, using proxy‑guided learning with operation decomposition and region‑aware aesthetic ranking to bootstrap the model. A competence‑based dynamic scheduler shifts training focus from proxy tasks to the main editing task as skills improve, and all reasoning traces for fine‑tuning are self‑synthesized by the model itself. Experiments on MIT‑Adobe FiveK and a new professionally retouched benchmark, SPIRE, show PrismGPT achieves state‑of‑the‑art results using only about 6% of the training data required by previous methods.