arXiv AI

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

arXiv:2608. 08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI.

arXiv Machine Learning
Aug 27

Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning

The paper introduces a novel framework that combines vision‑language model (VLM) generated preferences with the Plackett‑Luce (PL) ranking model for reward learning in reinforcement learning. Unlike traditional pairwise Bradley‑Terry approaches, the PL formulation allows listwise rankings of multiple candidates, enabling the use of different ranking sizes (K = 3, 4, 5). Experiments on Meta‑World manipulation tasks show that PL‑based reward models train robotic policies as effectively as, or better than, pairwise, K‑wise, and RL‑VLM‑F baselines, achieving up to an 86% mean final success rate and matching the Oracle baseline on the Drawer Open task.

By Srivalli Katkuri, Maxwell Kawada, Juan Wachs
arXiv AI
Jul 24

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.

By Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J. Ratliff, Jiafei Duan, Dieter Fox, Ranjay Krishna
arXiv Machine Learning
Aug 11

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

arXiv:2608. 09853v1 Announce Type: cross Abstract: General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored.

By Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, Zhian Su, Hang Guo, Tong Lu, Zhaofeng Xu, Jiahao Tang, Jianfei Yang, Donglin Wang, Peixi Peng, Mingxiu Chen, Deli Zhao, Xin Li
arXiv AI
Aug 7

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

arXiv:2607. 28609v2 Announce Type: replace Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world.

By Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
arXiv Computer Vision
Sep 4

WorldReward: Reward Modeling for Camera-Conditioned World Models

WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.

By Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
arXiv Machine Learning
Jun 29

Qwen-Image-2.0-RL Technical Report

arXiv:2606. 27608v1 Announce Type: cross Abstract: We present Qwen-Image-2.

By Yixian Xu, Kaiyuan Gao, Yuxiang Chen, Yilei Chen, Zecheng Tang, Zihao Liu, Zikai Zhou, Deqing Li, Hao Meng, Kuan Cao, Jiahao Li, Jie Zhang, Liang Peng, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiaoyue Chen, Yan Shu, Yanran Zhang, Yi Wang, Yu Wu, Yujia Wu, Zekai Zhang, Zhendong Wang, Xiao Xu, Kun Yan, Chenfei Wu
arXiv Machine Learning
Aug 19

Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

Prism‑GRPO enhances the GRPO reinforcement‑learning algorithm for vision‑language‑action policies by adding a weighted trajectory‑level execution‑quality score to binary success rewards. This approach splits groups with identical outcomes into a quality spectrum, preserving training signal while ensuring successes always outrank failures. Experiments on four RoboTwin tasks show Prism‑GRPO achieves higher success and quality at matched rollout budgets, reaching target success rates with up to 56% fewer rollouts and mitigating reward‑hacking behaviors that transfer to real‑robot deployment.

By Zeyun Deng, Yuzhe Lu, Yawei Wang, Linbo Liu, Qing Ping, Han Ding, Guande Wu, Panpan Xu, Jun Huan
arXiv Machine Learning
Jun 18

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.

By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal