arXiv AI By Jiayu Yao, Yiwei Wang, Anmeng Zhang, Zhe Sun, Songsong Wang, Lingrui Mei, Yuyao Ge, Shenghua Liu

Multimodal Reward Hacking in Reinforcement Learning

Read the original on arXiv AI →

arXiv:2607. 09492v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

Improving Generalization Robustness of Multimodal RLVR

arXiv:2608. 08802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA.

By Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You
arXiv Machine Learning
1d ago

Same Reward, Different Skills: When Multimodal RL Learns to Look

The paper demonstrates that reinforcement learning with verifiable rewards (RLVR) can improve vision‑language benchmark performance even when models are trained without visual input. When real images are introduced at test time, models trained blind recover about half of the performance gain at 3B parameters and nearly four‑fifths at 7B, but extended real‑image training can erode grounding while benchmark gains persist. The authors propose a visual resolvability rule and show that requiring visual evidence for correct answers leads to significant improvements in target discovery and generalization to unseen question types, while controls confirm that the gains stem from actual visual grounding rather than artifacts.

By Haocun Ye, Xinlong Jiang, Qile Chen, Bingyu Wang, Teng Zhang, Shubai Chen, Tingyu Wu, Zhenkun Zheng, Yiqiang Chen
Hugging Face Trending Papers
Aug 9

Improving Generalization Robustness of Multimodal RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two issues of the standard RL objective.