The paper introduces PIVOT, a dual-level learning framework designed to improve visually-grounded multimodal reasoning in large vision-language models. PIVOT employs a self‑calibrated experience replay mechanism to selectively reuse valuable visual reasoning trajectories, and a vision‑guided advantage allocation scheme that assigns extra rewards to tokens with strong visual support. Experiments on multiple benchmarks show that PIVOT enhances the multimodal reasoning performance of these models.
By Xinxin Song, Siyuan Li, Tingxiong Xiao, Jinli Suo
The paper introduces Activation Replay, a training‑free method that improves reasoning in post‑trained large multimodal models (LMMs) by replaying low‑entropy activations from the base model’s input context. It shows that Reinforcement Learning with Verifiable Rewards (RLVR) shifts low‑entropy activations and that modulating these activations enhances reasoning across tasks such as mathematics, visual agents, and video reasoning. Experiments demonstrate that Activation Replay outperforms alternatives like high‑entropy replay or direct cross‑model intervention, boosting Pass@K and broadening RLVR’s reasoning coverage.
By Yun Xing, Xiaobin Hu, Qingdong He, Jiangning Zhang, Shuicheng Yan, Shijian Lu, Yu-Gang Jiang
arXiv:2607. 14682v1 Announce Type: new Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge.
By Harikrishnan P M, Goutham Vignesh, Ganesh Parab, Saisubramaniam Gopalakrishnan, Vishal Vaddina, Varun V, Rohit Agrawal
The paper introduces VIG (Visual Information Gain), an information‑theoretic reward that evaluates each token in a multimodal chain‑of‑thought by measuring how much the image reduces its predictive uncertainty. VIG is computed online using two forward passes—one with and one without the image—eliminating the need for reference chains or external annotations. Experiments on six multimodal reasoning benchmarks and multiple Qwen3‑VL‑Thinking model sizes show that VIG consistently improves the accuracy–efficiency trade‑off, demonstrating that efficient multimodal reasoning arises from increasing visual information density rather than merely limiting chain length.
By Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang
arXiv:2609.21675v1 Announce Type: new
Abstract: Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural...
By Wan Xu, Yuanfan Guo, Kevin Han, LaLa Chen, Wangmeng Zuo
arXiv:2608.22429v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this relian...
By Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie
The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge
arXiv:2606. 07000v1 Announce Type: new Abstract: Recent post-training methods, particularly Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced the reasoning ability of Large Vision-Language Models (LVLMs).
By Shizhe Xiang, Ke An, Wenlong Yu, Yue Liu, Jian Luan, Pei Fu, Qilong Wang
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
By Peng, Lee, Yin Zhang, Yanglin Zhang, Haonan Wu, Zishan Liu, Ruoxi Zang, Xin Zhu, Jiayin Zheng, Jian Yao, Zefeng Ji, Fei Ma
arXiv:2606. 03937v1 Announce Type: new Abstract: While token-level entropy is commonly recognized as effective for credit assignment in text-only reinforcement learning with verifiable rewards (RLVR), it remains unclear whether this mechanism still holds in visual reasoning.
By Senjie Jin, Peixin Wang, Boyang Liu, Xiaoran Fan, Shuo Li, Zhiheng Xi, Jiazheng Zhang, Yuhao Zhou, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2606. 17888v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxiliary signals, failing to capture the intricate and sample-specific dependencies between text and images in mathematical problem-solving.
By Wanshi Xu, Haokun Zhao, Haidong Yuan, Songjun Cao, Long Ma
The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.
By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan