VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.
By Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.
By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson
The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning.
"whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."
By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
arXiv:2603. 23867v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts.
By Weixin Chen, Antonio Vergari, Han Zhao
arXiv:2604. 04917v3 Announce Type: replace-cross Abstract: What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks?
By Gabriel Sarch, Linrong Cai, Qunzhong Wang, Haoyang Wu, Danqi Chen, Zhuang Liu
UniCAR‑RL is an annotation‑free reinforcement learning framework designed to improve multimodal large language models’ visual mathematics reasoning. It decouples perception and reasoning by using three branches: Caption‑RL for perception optimization, Reasoning‑RL for logical reasoning with a gold image description, and QA‑RL for end‑to‑end question answering. Experiments show significant gains in mathematical and visual reasoning across various model architectures and scales using only raw short‑answer data.
By Yuzhe Li, Hao Yan, Hao Wang, Xingchen Liu, Ya-Qi Yu, Jihao Wu, Minghui Liao, Wei Chen, Yuliang Liu
arXiv:2606. 26196v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence.
By Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv
arXiv:2605. 14054v2 Announce Type: replace Abstract: Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs).
By Haozhe Wang, Qixin Xu, Changpeng Wang, Taofeng Xue, Chong Peng, Wenhu Chen, Fangzhen Lin
arXiv:2610.02117v1 Announce Type: cross
Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a froze...
By Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris
arXiv:2601. 10922v2 Announce Type: replace Abstract: We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constant and the main degree of freedom is the training data.
By Yosub Shin, Michael Buriek, Boris Sobolev, Pavel Bushuyeu, Vikas Kumar, Haoyang Xu, Samuel Watson, Igor Molybog
The paper introduces PCSR-Bench, a benchmark of 84,373 question‑answer pairs derived from 2,600 omnidirectional images across 26 indoor environments, designed to evaluate perspective‑conditioned spatial reasoning (PCSR) in multimodal large language models (MLLMs). It reports a significant perception–reasoning gap, with accuracy dropping from 57.59% on limited field‑of‑view reasoning to as low as 0.64% on open‑ended compositional directional chains. An RL‑based diagnostic study on a 7B‑scale model shows that reward shaping can improve performance to 60.06% on a controlled task, indicating partial plasticity of PCSR capabilities.
By Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras, Xu Zheng
arXiv:2609.31456v1 Announce Type: new
Abstract: Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hyp...
By Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy