arXiv:2608. 14430v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards.
By Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He
The paper investigates how different estimators of the reverse Kullback–Leibler (KL) divergence used as a regularization term in reinforcement learning (RL) training of large language models (LLMs) affect training stability and downstream performance. By analyzing gradient bias across various estimator configurations, the authors demonstrate that biased gradients can cause training instabilities, while unbiased configurations improve performance on both in‑domain and out‑of‑domain tasks. Experiments on Qwen2.5‑7B, Llama‑3.1‑8B‑Instruct, and Qwen3‑4B‑Instruct‑2507 confirm these findings and show that KL regularization also stabilizes off‑policy RL training in asynchronous setups.
By Vedant Shah, Johan Obando-Ceron, Vineet Jain, Brian Bartoldson, Bhavya Kailkhura, Sarthak Mittal, Glen Berseth, Pablo Samuel Castro, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain, Siddarth Venkatraman, Aaron Courville
arXiv:2608.23664v1 Announce Type: cross
Abstract: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing method...
By Jaemoo Choi, Wei Guo, Yuchen Zhu, Arash Vahdat, Molei Tao, Julius Berner, Yongxin Chen
arXiv:2601. 00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference.
By Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan
arXiv:2607. 07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.
By Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
arXiv:2606. 02884v1 Announce Type: cross Abstract: Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time.
By Sanjit Dandapanthula, Nicholas M. Boffi