arXiv AI

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

arXiv:2608. 03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL).

arXiv Machine Learning
Jul 21

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

arXiv:2509. 02522v3 Announce Type: replace-cross Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have empowered large language models (LLMs) to tackle challenging reasoning tasks such as mathematics and programming, however existing RLVR methods often suffer from sparse reward signals and unstable policy gradient updates inherent to RL-based approaches.

By Jiaming Li, Longze Chen, Ze Gong, Yukun Chen, Lu Wang, Wanwei He, Run Luo, Min Yang
arXiv AI
Jun 2

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

arXiv:2602. 10623v2 Announce Type: replace-cross Abstract: Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedback, yet they are often vulnerable to reward hacking due to noisy annotations and systematic biases such as response length or style.

By Zhibin Duan, Guowei Rong, Zhuo Li, Bo Chen, Mingyuan Zhou, Dandan Guo