arXiv Machine Learning

Semantic Editing with Coupled Stochastic Differential Equations

arXiv:2509. 24223v2 Announce Type: replace Abstract: Editing the content of an image with a pretrained text-to-image model remains challenging.

arXiv Computer Vision
Sep 11

Overpainting: Localized Context-aware Diffusion Image Editing

The paper introduces "overpainting," a localized, context-aware image editing technique that allows users to specify precise or loose editing regions via a trimap. The method adapts a pretrained diffusion model with joint attention and low‑rank adaptation, incorporating attention‑dropout to balance noise, source, and mask inputs. An automated pipeline generates training data by pairing images from language‑based editing models, curating them, and extracting trimaps, enabling the model to perform a wide range of editing tasks.

By Sam Sartor, Iliyan Georgiev, Michael Fischer, Valentin Deschaintre, Pieter Peers
arXiv AI
Jul 1

Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models

arXiv:2603. 12893v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment.

By David McAllister, Miika Aittala, Tero Karras, Janne Hellsten, Angjoo Kanazawa, Timo Aila, Samuli Laine
arXiv AI
Sep 16

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.

By Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra
arXiv Computer Vision
Sep 14

Unified Text-Image Generation with Weakness-Targeted Post-Training

The paper introduces a post‑training approach that enables a single inference process to transition from text reasoning to image synthesis, eliminating the need for explicit modality switching. Using the 14B BAGEL model, the authors demonstrate that targeted post‑training data and reward‑weighted training improve multimodal image generation across four independent T2I benchmarks. The study highlights the benefits of joint text‑image generation and strategic data selection for enhancing T2I performance.

By Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han, Yushi Hu, Emily Dinan, Amita Kamath, Michal Drozdzal, Reyhane Askari-Hemmat, Luke Zettlemoyer, Marjan Ghazvininejad