The paper introduces Rita, a reinforcement learning framework that addresses thinking drift in vision‑language models by enforcing consistency between reasoning and answers. Rita employs two reasoning‑label‑free rewards—thinking and consistency rewards—derived from the conditional probability of reference answers, and uses a difficulty‑aware data filtering strategy to select informative samples for training. Experiments on EgoIntention and RefEgo‑Int benchmarks demonstrate that Rita outperforms both supervised fine‑tuning and vanilla RL‑fine‑tuned approaches.
By Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao
arXiv:2610.02015v1 Announce Type: cross
Abstract: Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)-...
By Michael Sullivan, Alexander Koller
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
By Peng, Lee, Yin Zhang, Yanglin Zhang, Haonan Wu, Zishan Liu, Ruoxi Zang, Xin Zhu, Jiayin Zheng, Jian Yao, Zefeng Ji, Fei Ma
The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
By Jingpei Wu, Xiao Han, Weixiang Shen, Boer Zhang, Zifeng Ding, Volker Tresp
arXiv:2608. 03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL).
By Pyrros Koussios, Chenhao Li, Xin Chen, Andreas Krause
arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.
By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi
arXiv:2507.21931v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks....
By Carel van Niekerk, Renato Vukovic, Benjamin Ruppik, Hsien-chin Lin, Shutong Feng, Milica Ga\v{s}i\'c
arXiv:2603. 16728v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly deployed in high-stakes settings where reliable uncertainty quantification (UQ) is as important as predictive accuracy.
By Robert Welch, Emir Konuk, Kevin Smith
arXiv:2606. 16122v1 Announce Type: new Abstract: Visual thinking should not only sound right; it should show its evidence.
By Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang
arXiv:2605. 14054v2 Announce Type: replace Abstract: Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs).
By Haozhe Wang, Qixin Xu, Changpeng Wang, Taofeng Xue, Chong Peng, Wenhu Chen, Fangzhen Lin
LineupRL introduces a reinforcement learning framework with verifiable rewards for time series captioning, using a frozen large language model to identify the correct time series from a set of distractors based on a generated caption. This approach bypasses the limitations of supervised fine‑tuning and traditional RL rewards that poorly transfer to open‑ended time series generation. Experiments on two captioning benchmarks, as well as forecasting and reconstruction tasks, show that LineupRL outperforms both SFT and RL baselines across all metrics, and its trained 3B vision‑language model surpasses a 72B model distilled from SFT captions. The method also demonstrates resistance to reward hacking and produces captions that accurately trace trends and name key values.
By Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen