arXiv:2608. 06310v1 Announce Type: new Abstract: Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models.
By Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang, Tongran Liu, Jingbo Zhu
The paper introduces a novel framework that combines vision‑language model (VLM) generated preferences with the Plackett‑Luce (PL) ranking model for reward learning in reinforcement learning. Unlike traditional pairwise Bradley‑Terry approaches, the PL formulation allows listwise rankings of multiple candidates, enabling the use of different ranking sizes (K = 3, 4, 5). Experiments on Meta‑World manipulation tasks show that PL‑based reward models train robotic policies as effectively as, or better than, pairwise, K‑wise, and RL‑VLM‑F baselines, achieving up to an 86% mean final success rate and matching the Oracle baseline on the Drawer Open task.
By Srivalli Katkuri, Maxwell Kawada, Juan Wachs
arXiv:2608. 08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI.
By Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang
arXiv:2605.11151v3 Announce Type: replace
Abstract: Offline-to-online reinforcement learning (RL) improves sample efficiency by leveraging pre-collected datasets prior to online interaction. A key ch...
By Andrew Choi, Wei Xu
arXiv:2607. 01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards.
By Yuriy Maksyuta, George Bredis, Ruslan Rakhimov, Daniil Gavrilov
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by...