The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.
By Jaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk, Bo Yuan, Jinbin Bai, Yi Xin, Molei Tao, Yongxin Chen
arXiv:2609.37227v1 Announce Type: new
Abstract: Inference-time steering adapts pretrained diffusion and flow-based models to new tasks, e.g., to generate samples from a conditional distribution or sa...
By Adhithyan Kalaivanan, Zheng Zhao, Jens Sj\"olund, Fredrik Lindsten
arXiv:2609.31882v2 Announce Type: replace-cross
Abstract: Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to esti...
By Zhengyi Guo, Jiayuan Sheng, Wenpin Tang, David D. Yao
Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many applications it is desirable to steer according to distribution-level rewards, for example to calibrate with population-level information or to encourage diversity.
arXiv:2608. 08770v1 Announce Type: cross Abstract: Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function.
By Samuel Howard, Nikolas N\"usken
The paper introduces a data‑free learning objective called relative trajectory balance for training diffusion models to sample from a posterior defined by a diffusion prior and an arbitrary black‑box constraint or likelihood. It proves asymptotic correctness of this objective and demonstrates its use across vision, language, and multimodal tasks, including classifier guidance, language infilling, and text‑to‑image generation. Additionally, the method is applied to continuous control with a score‑based behavior prior, achieving state‑of‑the‑art results in offline reinforcement learning.
By Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Esmeralda S. Whitammer