The paper introduces Causal Visual Recurrent Reasoning (CVRR), a method that forces multimodal models to rely on latent visual states by using recurrent computation as the sole image‑conditioned path to prediction. CVRR initializes recurrence from the question hidden state after a pretrained vision‑language model has processed the image, repeatedly updates this state while re‑reading the same visual evidence, and removes all other visual traces before decoding. Experiments on several benchmarks show that CVRR maintains strong performance while other latent reasoners lose visual competence, and causal interventions confirm that predictions depend on the recurrent visual trajectory.
By Suhyeong Park, Junha Jung, Jaewoo Kang
The paper investigates latent visual reasoning in multimodal large language models, treating input, latent tokens, and final answer as a causal chain. Causal mediation analysis reveals two disconnections: latent tokens largely ignore input perturbations, and changes to latent tokens minimally affect the final answer, indicating limited causal influence. Probing shows latent tokens encode little visual information and are highly similar, leading the authors to propose CapImagine, an explicit text‑based imagination approach that outperforms latent‑space baselines on vision‑centric benchmarks.
By You Li, Chi Chen, Yanghao Li, Fanhu Zeng, Kaiyu Huang, Jinan Xu, Maosong Sun
arXiv:2610.01917v1 Announce Type: cross
Abstract: Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual...
By Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin
arXiv:2605.18641v2 Announce Type: replace
Abstract: Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generat...
By Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang, Saeed Vahidian, Wei-Lun Chao, Tanya Berger-Wolf, Yu Su, Raju Vatsavai, Jianyang Gu
The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.
By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan
arXiv:2603. 25629v2 Announce Type: replace-cross Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs).
By Andr\'e G. Viveiros, Nuno Gon\c{c}alves, Matthias Lindemann, Andr\'e Martins
Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses alignment between these latents and their visual targets, i.
arXiv:2511. 19418v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual perception, e.
By Yiming Qin, Bomin Wei, Jiaxin Ge, Konstantinos Kallidromitis, Stephanie Fu, Trevor Darrell, XuDong Wang
arXiv:2602. 00462v4 Announce Type: replace-cross Abstract: Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM.
By Benno Krojer, Shravan Nayak, Oscar Ma\~nas, Vaibhav Adlakha, Desmond Elliott, Siva Reddy, Marius Mosbach
arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.
By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.
By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi
arXiv:2604. 14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear.
By Danae S\'anchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras, Desmond Elliott