The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning.
"whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."
By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
UniCAR‑RL is an annotation‑free reinforcement learning framework designed to improve multimodal large language models’ visual mathematics reasoning. It decouples perception and reasoning by using three branches: Caption‑RL for perception optimization, Reasoning‑RL for logical reasoning with a gold image description, and QA‑RL for end‑to‑end question answering. Experiments show significant gains in mathematical and visual reasoning across various model architectures and scales using only raw short‑answer data.
By Yuzhe Li, Hao Yan, Hao Wang, Xingchen Liu, Ya-Qi Yu, Jihao Wu, Minghui Liao, Wei Chen, Yuliang Liu
arXiv:2512. 11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space.
By Chenrui Fan, Yijun Liang, Shweta Bhardwaj, Kwesi Cobbina, Ming Li, Tianyi Zhou
arXiv:2606. 12830v1 Announce Type: cross Abstract: While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction.
By Changye Li, Meng Lu, Yi Wu, Ligeng Zhu
arXiv:2511. 17731v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs).
By Lingxiao Li, Yifan Wang, Xinyan Gao, Chen Tang, Xiangyu Yue, Chenyu You
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
By Peng, Lee, Yin Zhang, Yanglin Zhang, Haonan Wu, Zishan Liu, Ruoxi Zang, Xin Zhu, Jiayin Zheng, Jian Yao, Zefeng Ji, Fei Ma