arXiv AI By Xingjian Tao, Yiwei Wang, Yujun Cai, Jing Tang

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

Read the original on arXiv AI →

arXiv:2607. 17243v1 Announce Type: new Abstract: Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning tends to grow verbose without becoming more accurate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 17

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.

By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
arXiv AI
Sep 17

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

The paper introduces PIVOT, a dual-level learning framework designed to improve visually-grounded multimodal reasoning in large vision-language models. PIVOT employs a self‑calibrated experience replay mechanism to selectively reuse valuable visual reasoning trajectories, and a vision‑guided advantage allocation scheme that assigns extra rewards to tokens with strong visual support. Experiments on multiple benchmarks show that PIVOT enhances the multimodal reasoning performance of these models.

By Xinxin Song, Siyuan Li, Tingxiong Xiao, Jinli Suo