arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.
By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal
arXiv:2609.36628v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, partic...
By Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe, Zhaohe Liao
arXiv:2608. 18607v2 Announce Type: replace Abstract: Using reinforcement learning to post-train joint video-audio generation models requires a reward signal.
By Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Xu Hang, Yu-Gang Jiang, Zuxuan Wu
arXiv:2608. 08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI.
By Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang
arXiv:2607. 15740v1 Announce Type: cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI.
By Bo-An Chang, Yu-Chih Chen
WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.
By Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
arXiv:2608.29804v1 Announce Type: new
Abstract: Virtual try-on (VTON) requires not only realistic generation but also faithful preservation of garment characteristics. However, existing evaluation me...
By Kaidong Zhang, Yukang Ding, Xiaoyu Liu, Ying Chen
PreResQ‑R1 introduces a Preference‑Response Disentangled Reinforcement Learning framework for Visual Quality Assessment that jointly optimizes absolute score regression and relative ranking consistency. It employs a dual‑branch reward system—modeling intra‑sample response coherence and inter‑sample preference alignment—trained with Group Relative Policy Optimization. The method extends to video quality assessment via a global‑temporal and local‑spatial data flow strategy, achieving state‑of‑the‑art results on 10 IQA and 5 VQA benchmarks with only 6K images and 28K videos, and provides human‑aligned reasoning traces.
By Zehui Feng, Weichuan Wang, Xiaohan Chen, Ting Han
arXiv:2608.21425v1 Announce Type: cross
Abstract: Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating genera...
By Nai-Xin Zhai, Weihua Cheng, Dexu Yu, Yikai Gu, Hanwen Du, Junchen Fu, Chenxi Huang, Yingwei Song, Liyuan Lillian Ma, Yang Ran, Youhua Li, Yongxin Ni
We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-Be...
The paper introduces a novel framework that combines vision‑language model (VLM) generated preferences with the Plackett‑Luce (PL) ranking model for reward learning in reinforcement learning. Unlike traditional pairwise Bradley‑Terry approaches, the PL formulation allows listwise rankings of multiple candidates, enabling the use of different ranking sizes (K = 3, 4, 5). Experiments on Meta‑World manipulation tasks show that PL‑based reward models train robotic policies as effectively as, or better than, pairwise, K‑wise, and RL‑VLM‑F baselines, achieving up to an 86% mean final success rate and matching the Oracle baseline on the Drawer Open task.
By Srivalli Katkuri, Maxwell Kawada, Juan Wachs
arXiv:2606. 27180v1 Announce Type: cross Abstract: Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory.
By Henrik M\"uller, Daniel Kudenko