The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.
By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv:2608.23664v1 Announce Type: cross
Abstract: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing method...
By Jaemoo Choi, Wei Guo, Yuchen Zhu, Arash Vahdat, Molei Tao, Julius Berner, Yongxin Chen
arXiv:2601. 00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference.
By Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan
arXiv:2607. 07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.
By Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
arXiv:2512. 14617v2 Announce Type: replace-cross Abstract: Many practical decision-making problems involve tasks whose success depends on the entire system history, rather than on achieving a state with desired properties.
By Alessandro Trapasso, Luca Iocchi, Fabio Patrizi
arXiv:2606. 15048v1 Announce Type: new Abstract: Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between predictions along the denoising trajectory.
By Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu, Junfeng Wen
arXiv:2510. 04019v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-training techniques, including reinforcement learning (RL), remains a key challenge for dLLMs, especially for downstream applications.
By Anthony Zhan
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences.
arXiv:2606. 27766v1 Announce Type: cross Abstract: Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe.
By Shiqiang Gong
arXiv:2610.00661v1 Announce Type: cross
Abstract: Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in paralle...
By Yue YU, Bowen Zuo, David Crandall, Yinglun Zhu, Dongruo Zhou
The paper introduces reinforcement learning for Continuous-Time Jump Markov Decision Processes (CTJMDPs) with general discrete state spaces and continuous/discrete actions. It develops entropy‑regularized continuous‑time control and establishes theoretical foundations for q‑learning in this setting, providing model‑free algorithms that outperform naive discretization. Numerical tests on network dynamic pricing demonstrate the method’s ability to learn near‑optimal policies and scale to large networks.
By Huiling Meng, Ningyuan Chen, Xuefeng Gao
The paper introduces dFlowGRPO, a reinforcement learning framework tailored for discrete flow models (DFMs). It generalizes previous work on diffusion large language models by supporting various probability paths and non-masked source distributions, and formulates denoising as a Markov decision process that leverages transition rates and posterior models. Experiments on the multimodal DFM FUDOKI show that dFlowGRPO outperforms existing GRPO methods on text‑to‑image generation and matches continuous flow models on multimodal understanding tasks.
By Zhengyan Wan, Yidong Ouyang, Panwen Hu, Qiang Sun