arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2608.29107v1 Announce Type: new
Abstract: While modern generative models excel at modeling complex data, precise inference-time control in conditional generation remains a critical challenge. C...
By Avishag Nevo, Tamir Hazan
arXiv:2606. 08602v1 Announce Type: cross Abstract: We present an online reinforcement learning (RL) algorithm for fine-tuning flow-matching policies in continuous-control problems.
By Boshu Lei, Kostas Daniilidis, Antonio Loquercio
arXiv:2609.15123v1 Announce Type: cross
Abstract: Flow-based policies offer an expressive representation for online reinforcement learning, but conventional flow matching requires samples drawn from...
By Bumgeun Park, Hyukjun Yang, Donghwan Lee
arXiv:2601. 23231v2 Announce Type: replace-cross Abstract: Flow-based generative models provide strong unconditional priors for inverse problems, but guiding their dynamics for conditional generation remains challenging.
By George Webber, Alexander Denker, Riccardo Barbano, Andrew J Reader
arXiv:2607. 21302v1 Announce Type: new Abstract: Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by leveraging policy priors derived from offline demonstrations.
By Gong Gao, Weidong Zhao, Xianhui Liu, Ning Jia
arXiv:2606. 29820v1 Announce Type: cross Abstract: In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging.
By Qijun Li, Zheng Fu, Qi Song, Yifei He, Weitao Zhou, Kun Jiang, Diange Yang
arXiv:2602. 21429v3 Announce Type: replace Abstract: Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distributions.
By Darshan Gadginmath, Ahmed Allibhoy, Fabio Pasqualetti
arXiv:2606. 30376v1 Announce Type: new Abstract: Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods.
By Zheming Fu, Ruizhe He, Wei Shang, Xiaoxiao Ma, Lei Wang, Chang Liu, Siming Fu
arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.
By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
The paper studies classifier‑free guidance (CFG) in Flow Matching, showing that strong guidance can distort the generated distribution by shifting the mean and concentrating trajectories. By interpreting Flow Matching as a time‑varying gradient flow, the authors explain how CFG reshapes the underlying potential and propose a training‑free method, Posterior‑Mean‑Capped CFG (PMC‑CFG), that adaptively limits guidance to the strongest feasible level. Experiments on synthetic and large‑scale image‑generation tasks demonstrate that PMC‑CFG reduces distortion and concentration while improving the alignment–diversity trade‑off, especially when nominal guidance is large.
By Jishen Peng, Zheng Ma
arXiv:2607. 10369v1 Announce Type: cross Abstract: Flow-matching policies have emerged as an effective policy parameterization for robot learning.
By Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao