arXiv AI By Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf

Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models

Read the original on arXiv AI →

The paper introduces MOSAIC, a benchmark that tests how vision‑language models translate theory‑of‑mind (ToM) reasoning into coordinated social actions across verbal and nonverbal channels. Across 200 trials per model, 13 models—including 11 VLMs—failed to align their behaviors with ToM constraints, revealing bottlenecks in generating coherent nonverbal signals and interpreting others’ actions. A structured model with an explicit ToM module, PCM‑LLM, succeeded in all conditions, indicating that belief‑action coupling is crucial for these tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 31

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

arXiv:2607. 26769v1 Announce Type: cross Abstract: Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states.

By Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
arXiv AI
Sep 1

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

arXiv:2605.30117v2 Announce Type: replace Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VL...

By Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang, Jiayu Hu, Haozhe Shan, Han Dong, Jinpeng Lu, Yinda Chen, Yi Zhang, Yong Dai, Xiaozhu Ju