arXiv:2609.36572v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further e...
By Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang
The paper introduces EASE, a method that enhances multimodal reinforcement learning with verifiable rewards (RLVR) by adding visual‑evidence process supervision. EASE transforms annotated evidence regions into smoothed visual‑token targets and uses them to guide attention during RL training, but only on high‑reward trajectories. Experiments on Qwen2.5‑VL‑7B, Qwen3‑VL‑4B, and Qwen3‑VL‑8B show that EASE improves average scores over DAPO by 2.5 to 3.1 points across perception, hallucination, visual math, and multimodal reasoning benchmarks, and diagnostics confirm better alignment of visual attention with annotated evidence.
By Ruina Hu, Chen Wang, Lai Wei, Jionghao Bai, Bin Yu, Weiran Huang, Kai Wang, Yue Wang
arXiv:2609. 31450v1 Announce Type: cross Abstract: Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy.
By Wang Jingxin
arXiv:2608.21595v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning ability of vision-language models (VLMs), and diversifying the rollouts wi...
By Michael Jerge, Joseph Pelczar, Justin Downes
arXiv:2607. 09492v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance.
By Jiayu Yao, Yiwei Wang, Anmeng Zhang, Zhe Sun, Songsong Wang, Lingrui Mei, Yuyao Ge, Shenghua Liu
SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.
By Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
The paper introduces Causal Visual Recurrent Reasoning (CVRR), a method that forces multimodal models to rely on latent visual states by using recurrent computation as the sole image‑conditioned path to prediction. CVRR initializes recurrence from the question hidden state after a pretrained vision‑language model has processed the image, repeatedly updates this state while re‑reading the same visual evidence, and removes all other visual traces before decoding. Experiments on several benchmarks show that CVRR maintains strong performance while other latent reasoners lose visual competence, and causal interventions confirm that predictions depend on the recurrent visual trajectory.
By Suhyeong Park, Junha Jung, Jaewoo Kang
The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.
By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that...
arXiv:2610.01973v1 Announce Type: new
Abstract: Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: som...
By Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu
arXiv:2609.39168v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing...
By Zhihan Zhang, Lizi Liao
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
By Peng, Lee, Yin Zhang, Yanglin Zhang, Haonan Wu, Zishan Liu, Ruoxi Zang, Xin Zhu, Jiayin Zheng, Jian Yao, Zefeng Ji, Fei Ma