Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences.
CAST introduces a reinforcement‑learning fine‑tuning framework for diffusion models that addresses three key limitations: it automatically selects the denoising window based on each model’s trajectory, decomposes prompts into verifiable semantic atoms via Causal Scene Graphs, and applies atom‑level rewards spatially weighted in the policy objective. The method is applied to FLUX.2‑dev and Qwen‑Image‑2512, yielding up to 3.07× improvement on the hardest GenEval 2 prompts compared with Flow‑GRPO while also enhancing overall generation quality.
By Shu Yu, Chaochao Lu
arXiv:2607. 07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.
By Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
arXiv:2608. 14430v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards.
By Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He
The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.
By Jaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk, Bo Yuan, Jinbin Bai, Yi Xin, Molei Tao, Yongxin Chen
arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.
By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
The paper introduces Reflection-Aware GRPO (RA‑GRPO), a reinforcement‑learning framework that aligns diffusion generative models with human preferences. It uses Diffusion Reflection to correct intermediate sampling paths by reversing the diffusion process, and Counterfactual Path Synthesis to embed these corrected trajectories into the policy, avoiding extra inference cost. Experiments on text‑to‑image and text‑to‑video models show RA‑GRPO outperforms existing methods, reducing reward hacking and improving generalization while remaining architecture‑agnostic.
By Junlong Wu, Jiuzhou Lin, Jia Sun, Boheng Zhang, Huaiqing Wang, Dewen Fan, Houde Liu, Qianqian Gan, Fan Yang, Tingting Gao
The paper introduces DiffusionOPSD, an on‑policy self‑distillation framework that transforms image‑level reinforcement learning rewards into explicit targets for intermediate denoising predictions in diffusion models. By generating trajectories with a frozen behavior policy and constructing bounded positive and negative targets around query states, the method trains a policy to fit these targets before updating the behavior policy via an exponential moving average. Experiments on SD 3.5‑M and Z‑Image‑Turbo show that DiffusionOPSD achieves the best held‑out scores in 19 of 20 reward‑matched settings, outperforms the strongest competitor by up to 44 % and cuts GPU‑hour usage by 40–63 % compared to DiffusionNFT.
By Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
arXiv:2606. 18066v1 Announce Type: new Abstract: We introduce the Noise-Tilted Reverse Kernel (NTRK), a reward-guided diffusion sampler that injects reward gradients through the noise term, leaving the pretrained reverse kernel unchanged and requiring only a single sample per step.
By Jisung Hwang, Yunhong Min, Jaihoon Kim, I-Chao Shen, Minhyuk Sung
arXiv:2609. 22041v1 Announce Type: new Abstract: Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback.
By Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa
Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for flow models therefore replace the inference-time ODE with a stochastic differential equation (SDE) during training.
The paper introduces Objective-aware Trajectory Credit Assignment (OTCA), a framework that refines reinforcement learning for diffusion-based visual generation. OTCA decomposes credit across denoising steps and allocates multiple reward signals adaptively, addressing the coarse, uniform reward assignment of existing GRPO pipelines. Experiments demonstrate that OTCA consistently enhances image and video generation quality across various metrics.
By Rui Li, Ke Hao, Yuanzhi Liang, Haibin Huang, Chi Zhang, Yun Gu, Xuelong Li