arXiv:2608. 08770v1 Announce Type: cross Abstract: Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function.
By Samuel Howard, Nikolas N\"usken
arXiv:2606. 02884v1 Announce Type: cross Abstract: Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time.
By Sanjit Dandapanthula, Nicholas M. Boffi
arXiv:2606. 13240v1 Announce Type: cross Abstract: A key strength of diffusion models lies in their flexibility, since their outputs can be controlled at sampling time through guidance.
By Rapha\"el Razafindralambo, R\'emy Sun, Fr\'ed\'eric Precioso, Jes Frellsen, Pierre-Alexandre Mattei
arXiv:2607. 07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.
By Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
arXiv:2602. 03211v2 Announce Type: replace-cross Abstract: Diffusion models have demonstrated strong generative performance; however, generated samples often fail to fully align with human intent.
By Yeongmin Kim, Donghyeok Shin, Byeonghu Na, Minsang Park, Richard Lee Kim, Il-Chul Moon
arXiv:2507. 08390v5 Announce Type: replace Abstract: Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-scale training.
By Meihua Dang, Jiaqi Han, Minkai Xu, Kai Xu, Akash Srivastava, Stefano Ermon
arXiv:2502. 04646v2 Announce Type: replace-cross Abstract: Weighted sampling -- sampling from a probability density function (PDF) proportional to the product of a base PDF and a weight function -- is a fundamental technique with wide-ranging applications in variance reduction, biased sampling, data augmentation, and more.
By Heasung Kim, Taekyun Lee, Hyeji Kim, Gustavo de Veciana
arXiv:2607. 01144v1 Announce Type: cross Abstract: While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow regions of the underlying distribution.
By Binglin Ji, Anindya Sarkar, Hengchang Lu, Jens Sj\"olund, Yevgeniy Vorobeychik
arXiv:2606. 15359v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for planning and control by learning multimodal distributions over actions and trajectories.
By Paolo Giaretta, Zeyang Li, Navid Azizan
arXiv:2607. 14522v1 Announce Type: new Abstract: We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC).
By Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang
arXiv:2511. 06239v2 Announce Type: replace-cross Abstract: Learning-based methods for sampling from the Gibbs distribution in finite-dimensional spaces have progressed quickly, yet theory and algorithmic design for infinite-dimensional function spaces remain limited.
By Byoungwoo Park, Juho Lee, Guan-Horng Liu
arXiv:2608. 14430v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards.
By Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He