arXiv AI By Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

Read the original on arXiv AI →

The paper investigates why Vision Language Models (VLMs) struggle with abstract reasoning tasks. Using the Relational Match-to-Sample paradigm and mechanistic analysis, the authors identify four factors—capability tier, model scale, number of objects per scene, and lack of per-object noise—that influence VLM performance. They uncover two competing neural circuits: an early surface-feature circuit and a late relational circuit, and show that ablating relational heads degrades performance more than random heads, suggesting the relational circuit is broadly used.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 5

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

arXiv:2607. 27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning.

By Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang, Chen Zhang, Shunian Chen, Caijun Yan, Jianyao Xu, Shunyu Liu, Weijie Fu, Peiliang Li, Xiaozhi Chen, Yuxiang Cai
Hugging Face Trending Papers
Jul 14

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.

arXiv Machine Learning
Aug 27

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

VBVR-Pro is a closed‑loop testbed that enables native visual reasoning through generation, offering 300 procedurally generated tasks that scale training and allow strong transfer to external benchmarks. It supplies verifiable reward scorers based on deterministic, task‑specific rules, outperforming VLM‑as‑a‑judge approaches and providing reliable signals for reinforcement learning. The suite also facilitates controlled modality studies, revealing that video generation excels at persistent spatiotemporal tracking while interleaved generation offers a compute‑efficient alternative, and highlights the importance of vision‑native trajectories for reasoning.

By Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Rapha\"el Milli\`ere, Vincent C. M\"uller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai