Hugging Face Trending Papers

Improving Generalization Robustness of Multimodal RLVR

Read the original on Hugging Face Trending Papers →

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two issues of the standard RL objective.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 11

Improving Generalization Robustness of Multimodal RLVR

arXiv:2608. 08802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA.

By Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You
arXiv AI
Jun 19

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.

By Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei
arXiv Computer Vision
Aug 27

Boosting Reasoning in Large Multimodal Models via Activation Replay

The paper introduces Activation Replay, a training‑free method that improves reasoning in post‑trained large multimodal models (LMMs) by replaying low‑entropy activations from the base model’s input context. It shows that Reinforcement Learning with Verifiable Rewards (RLVR) shifts low‑entropy activations and that modulating these activations enhances reasoning across tasks such as mathematics, visual agents, and video reasoning. Experiments demonstrate that Activation Replay outperforms alternatives like high‑entropy replay or direct cross‑model intervention, boosting Pass@K and broadening RLVR’s reasoning coverage.

By Yun Xing, Xiaobin Hu, Qingdong He, Jiangning Zhang, Shuicheng Yan, Shijian Lu, Yu-Gang Jiang