The paper introduces Amortized Inpainting with Diffusion (AID), a method that keeps a pretrained diffusion backbone fixed and trains a small reusable guidance module offline for image inpainting. AID formulates the problem as deterministic guidance with a supervised terminal objective, derives an auxiliary Gaussian formulation to make it learnable, and proves that solving the randomized problem recovers the optimal deterministic guidance field. Experiments on AFHQv2, FFHQ, and ImageNet show that AID improves the quality–speed trade‑off over strong baselines while adding less than one percent trainable overhead.
By Yilie Huang, Xun Yu Zhou
The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.
By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU
arXiv:2606. 17979v1 Announce Type: new Abstract: Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory.
By Jinjie Shen, Wei Deng, Xian Hu, Daiguo Zhou, Jian Luan
CAST introduces a reinforcement‑learning fine‑tuning framework for diffusion models that addresses three key limitations: it automatically selects the denoising window based on each model’s trajectory, decomposes prompts into verifiable semantic atoms via Causal Scene Graphs, and applies atom‑level rewards spatially weighted in the policy objective. The method is applied to FLUX.2‑dev and Qwen‑Image‑2512, yielding up to 3.07× improvement on the hardest GenEval 2 prompts compared with Flow‑GRPO while also enhancing overall generation quality.
By Shu Yu, Chaochao Lu
The paper introduces DiffusionOPSD, an on‑policy self‑distillation framework that transforms image‑level reinforcement learning rewards into explicit targets for intermediate denoising predictions in diffusion models. By generating trajectories with a frozen behavior policy and constructing bounded positive and negative targets around query states, the method trains a policy to fit these targets before updating the behavior policy via an exponential moving average. Experiments on SD 3.5‑M and Z‑Image‑Turbo show that DiffusionOPSD achieves the best held‑out scores in 19 of 20 reward‑matched settings, outperforms the strongest competitor by up to 44 % and cuts GPU‑hour usage by 40–63 % compared to DiffusionNFT.
By Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
arXiv:2603.23463v2 Announce Type: replace-cross
Abstract: Recent diffusion-based models achieve photorealism in image inpainting but require many sampling steps, limiting practical use. Few-step text...
By Duc Vu, Kien Nguyen, Trong-Tung Nguyen, Ngan Nguyen, Phong Nguyen, Khoi Nguyen, Cuong Pham, Anh Tran
AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.
By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen
The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
Text-to-image (T2I) diffusion models have achieved striking progress but still struggle to synthesize rare concepts involving unusual attribute-object pairings, often resulting in concept omission or semantic drift where a dominant entity overwhelms the generation. Tracing these failures to a lack of compositional balance during the denoising trajectory, we propose RADIANCE, a training-free framework that treats inference as a closed-loop feedback process.
arXiv:2604. 26985v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) generate discrete sequences by iterative denoising under an absorbing masking process.
By Michael Cardei, Huu Binh Ta, Ferdinando Fioretto
The paper introduces a structured‑prior‑guided diffusion inpainting framework that enhances traffic sign augmentation by incorporating semantic, appearance, and geometric priors through text prompts, vector templates, and ControlNet. It enforces physical consistency with colour and edge‑structure losses, achieving superior reconstruction fidelity and OCR accuracy compared to larger models. The synthetic data generated improves detection performance for rare traffic sign classes by up to 7.4× over a real‑data‑only baseline.
By Luo Li, Chongchong Huang, Jun Jia, Qiang Gao, Xinlong Liu, Gui Yang, Liang Cao
arXiv:2608. 03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts.
By Ning Zhu, An Chen, Mengfei Zhao, Juntao Xu, Jingze Liang, Boyuan Gu, Liang-Jian Deng