Towards Robust Sequential Decomposition for Complex Image Editing
arXiv:2605. 09233v2 Announce Type: replace-cross Abstract: Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions.
arXiv:2510. 08532v2 Announce Type: replace-cross Abstract: Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language.
arXiv:2605. 09233v2 Announce Type: replace-cross Abstract: Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions.
arXiv:2607.18227v2 Announce Type: replace Abstract: In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and imag...
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge this gap, we present Goku, a large-scale dataset featuring 2 million high-quality, instruction-aligned video editing pairs, which is the first to extend task boundaries from basic appearance editing to multi-task and structural manipulations(e.
arXiv:2606. 08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes.
arXiv:2606. 19103v1 Announce Type: cross Abstract: Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions.
arXiv:2608.17566v2 Announce Type: replace Abstract: The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing...
arXiv:2606. 05950v1 Announce Type: new Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models.
CoinVE-200K is a large, high‑quality dataset for compositional instruction‑guided video editing, featuring 1080p video‑editing pairs up to 201 frames long and containing 2–5 atomic editing operations per sample. The dataset covers diverse editing intents—targeting humans, objects, and backgrounds with addition, removal, modification, and stylization—while ensuring instruction faithfulness, visual quality, temporal consistency, and compositional diversity through a careful generation and filtering pipeline. CoinVE-Bench benchmarks these capabilities, and CoinVE-Edit, a 22B model built on Wan2.1‑T2V‑14B and Qwen3‑VL‑8B‑Instruct, demonstrates strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs.
RefineEdit is a training‑free prompt‑to‑prompt image editing framework that uses a Generative Refinement Network to edit images by refining binary image codes. It couples edit localization with content generation, selecting editable positions based on signed probability differences between an editing branch and a source branch, and stabilizes edits with adaptive spatial freezing and finite bit locking. The method requires no additional training, external masks, or attention control, and outperforms other methods on PIE‑Bench in background‑preservation metrics and CLIP scores.
arXiv:2509. 24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data.
The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require considerable expertise to use effectively.