arXiv:2606. 08602v1 Announce Type: cross Abstract: We present an online reinforcement learning (RL) algorithm for fine-tuning flow-matching policies in continuous-control problems.
By Boshu Lei, Kostas Daniilidis, Antonio Loquercio
arXiv:2510. 09976v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models such as OpenVLA, Octo, and $\pi_0$ have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fundamentally constrained by the quality and coverage of supervised data.
By Mingyang Lyu, Yinqian Sun, Erliang Lin, Huangrui Li, Ruolin Chen, Feifei Zhao, Yi Zeng
arXiv:2602. 04879v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorithm.
By Penghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang, Chao Du, Min Lin, Wee Sun Lee
arXiv:2606. 09821v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a key component of post-training large language models (LLMs).
By Jiarui Yao, Xiangxin Zhou, Penghui Qi, Wee Sun Lee, Liefeng Bo, Tianyu Pang
arXiv:2609. 22041v1 Announce Type: new Abstract: Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback.
By Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa
arXiv:2606. 30376v1 Announce Type: new Abstract: Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods.
By Zheming Fu, Ruizhe He, Wei Shang, Xiaoxiao Ma, Lei Wang, Chang Liu, Siming Fu
arXiv:2603. 23086v2 Announce Type: replace Abstract: Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimization for sample quality and diversity.
By Orhun Bugra Baran, Melih Kandemir, Ramazan Gokberk Cinbis
arXiv:2609.38855v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic m...
By Gongxin Yao, Yongsheng Zhao, Jiayin Deng, Deng Liang, Han Gao, Lei Zhao, Baoping Cheng
arXiv:2512. 23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$.
By Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Qian Liu, Baoxiang Wang
arXiv:2607. 10369v1 Announce Type: cross Abstract: Flow-matching policies have emerged as an effective policy parameterization for robot learning.
By Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao
The paper introduces dFlowGRPO, a reinforcement learning framework tailored for discrete flow models (DFMs). It generalizes previous work on diffusion large language models by supporting various probability paths and non-masked source distributions, and formulates denoising as a Markov decision process that leverages transition rates and posterior models. Experiments on the multimodal DFM FUDOKI show that dFlowGRPO outperforms existing GRPO methods on text‑to‑image generation and matches continuous flow models on multimodal understanding tasks.
By Zhengyan Wan, Yidong Ouyang, Panwen Hu, Qiang Sun
The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.
By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao