arXiv Computer Vision By Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang

VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression

Read the original on arXiv Computer Vision →

The paper introduces VIG (Visual Information Gain), an information‑theoretic reward that evaluates each token in a multimodal chain‑of‑thought by measuring how much the image reduces its predictive uncertainty. VIG is computed online using two forward passes—one with and one without the image—eliminating the need for reference chains or external annotations. Experiments on six multimodal reasoning benchmarks and multiple Qwen3‑VL‑Thinking model sizes show that VIG consistently improves the accuracy–efficiency trade‑off, demonstrating that efficient multimodal reasoning arises from increasing visual information density rather than merely limiting chain length.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 14

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.

arXiv AI
Jul 17

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

arXiv:2607. 14682v1 Announce Type: new Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge.

By Harikrishnan P M, Goutham Vignesh, Ganesh Parab, Saisubramaniam Gopalakrishnan, Vishal Vaddina, Varun V, Rohit Agrawal
arXiv AI
Jul 15

Visual Access Boundaries in Vision-Language Model Reasoning

arXiv:2607. 12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces.

By Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
arXiv AI
Jun 2

Video Reasoning without Training

arXiv:2510. 17045v2 Announce Type: replace-cross Abstract: Video reasoning using Large Multimodal Models (LMMs) relies on costly reinforcement learning (RL) and verbose chain-of-thought, resulting in substantial computational overhead during both training and inference.

By Deepak Sridhar, Kartikeya Bhardwaj, Jeya Pradha Jeyaraj, Nuno Vasconcelos, Ankita Nayak, Harris Teague