The paper introduces DiffusionOPSD, an on‑policy self‑distillation framework that transforms image‑level reinforcement learning rewards into explicit targets for intermediate denoising predictions in diffusion models. By generating trajectories with a frozen behavior policy and constructing bounded positive and negative targets around query states, the method trains a policy to fit these targets before updating the behavior policy via an exponential moving average. Experiments on SD 3.5‑M and Z‑Image‑Turbo show that DiffusionOPSD achieves the best held‑out scores in 19 of 20 reward‑matched settings, outperforms the strongest competitor by up to 44 % and cuts GPU‑hour usage by 40–63 % compared to DiffusionNFT.
By Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We in...
arXiv:2512.22802v2 Announce Type: replace-cross
Abstract: Step distillation accelerates diffusion sampling by training a few-step student to imitate a many-step teacher, but distillation itself remai...
By Amirhossein Tighkhorshid, Zahra Dehghanian, Hamid R. Rabiee
arXiv:2609.30840v1 Announce Type: cross
Abstract: One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit ge...
By Austin Wang, Ziheng Cheng, Lexing Ying
arXiv:2609.38853v1 Announce Type: new
Abstract: Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their...
By Yifei Wang, Xiaoyu Wu, Tsu-Jui Fu, Chen Chen, Liang-Chieh Chen, Zhe Gan, Chen Wei
CAST introduces a reinforcement‑learning fine‑tuning framework for diffusion models that addresses three key limitations: it automatically selects the denoising window based on each model’s trajectory, decomposes prompts into verifiable semantic atoms via Causal Scene Graphs, and applies atom‑level rewards spatially weighted in the policy objective. The method is applied to FLUX.2‑dev and Qwen‑Image‑2512, yielding up to 3.07× improvement on the hardest GenEval 2 prompts compared with Flow‑GRPO while also enhancing overall generation quality.
By Shu Yu, Chaochao Lu