arXiv Computation and Language

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

arXiv AI
Aug 24

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

StateSight is a new benchmark designed to isolate and evaluate the ability of vision‑language models to reconstruct latent spatial structure from a single image. It consists of three procedurally generated task families—cube‑net opposite‑face reasoning, occluded cube‑tower counting, and 4‑neighbor connected‑component counting—each with 300 deterministic prompts and exact‑match scoring. The benchmark also includes a companion dataset, StateSight‑Steps, with 900 image‑text examples and 3,600 intermediate visual states to aid analysis of reconstruction errors.

By Michelle Lin
arXiv Computation and Language
Sep 1

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

arXiv:2606.01914v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly...

By Chuang Ma, Qianying Liu, Tomoyuki Obuchi, Fei Cheng, Wang Yang, Sudong Cai, Shuyuan Zheng, Akiko Aizawa, Sadao Kurohashi
arXiv Computer Vision
Aug 26

DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton

The paper introduces DoublesEval, a diagnostic framework that uses professional doubles badminton to test visual‑language models’ ability to reason about dynamic multi‑agent interactions. It decomposes rallies into key moments and evaluates models across four dimensions—atomic recognition, intra‑segment composite understanding, cross‑segment causal reasoning, and high‑level tactical abstraction—highlighting specific reasoning failures. The authors also propose TacticCheck, a lightweight consistency checker that improves performance without retraining the models, yet significant gaps remain in tactical reasoning.

By Jintao Cheng, Weibin Li
arXiv Computer Vision
5d ago

Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

The paper introduces “SpaceConflict”, a benchmark of 23,196 multimodal inputs designed to test whether large language models can not only report spatial facts but also use them in reasoning. Experiments show a gap between a model’s ability to recover an initial spatial state from visual evidence and its ability to apply that state to perform transformations, with the gap narrowing but not closing as model size increases. To address this, the authors propose Operational State Supervision (OSS), which supervises task‑relevant spatial states and their transformation trajectories, improving accuracy on judgments that require state organization and use.

By Jinchang Zhang, Guoyu Lu