The paper introduces Causal Visual Recurrent Reasoning (CVRR), a method that forces multimodal models to rely on latent visual states by using recurrent computation as the sole image‑conditioned path to prediction. CVRR initializes recurrence from the question hidden state after a pretrained vision‑language model has processed the image, repeatedly updates this state while re‑reading the same visual evidence, and removes all other visual traces before decoding. Experiments on several benchmarks show that CVRR maintains strong performance while other latent reasoners lose visual competence, and causal interventions confirm that predictions depend on the recurrent visual trajectory.
By Suhyeong Park, Junha Jung, Jaewoo Kang
The paper investigates latent visual reasoning in multimodal large language models, treating input, latent tokens, and final answer as a causal chain. Causal mediation analysis reveals two disconnections: latent tokens largely ignore input perturbations, and changes to latent tokens minimally affect the final answer, indicating limited causal influence. Probing shows latent tokens encode little visual information and are highly similar, leading the authors to propose CapImagine, an explicit text‑based imagination approach that outperforms latent‑space baselines on vision‑centric benchmarks.
By You Li, Chi Chen, Yanghao Li, Fanhu Zeng, Kaiyu Huang, Jinan Xu, Maosong Sun
arXiv:2610.01917v1 Announce Type: cross
Abstract: Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual...
By Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin
arXiv:2605.18641v2 Announce Type: replace
Abstract: Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generat...
By Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang, Saeed Vahidian, Wei-Lun Chao, Tanya Berger-Wolf, Yu Su, Raju Vatsavai, Jianyang Gu
The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.
By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan
arXiv:2603. 25629v2 Announce Type: replace-cross Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs).
By Andr\'e G. Viveiros, Nuno Gon\c{c}alves, Matthias Lindemann, Andr\'e Martins