CCRV-Bench is a constraint‑driven benchmark designed to evaluate visual causal reasoning in vision‑language models on single‑image physical scenarios. It assesses four causal task dimensions—causal relation discovery, state prediction, causal diagnosis, and intervention—while applying constraints such as entity symbolization, spatial grounding, factual adversarial constraints, and minimalist output constraints to reduce shortcut learning. Experiments on 15 multimodal models reveal that constraint sensitivity varies by task and model, with intervention and spatial grounding having the largest impact and factual adversarial constraints improving causal diagnosis across models.
By Linyuan Gao, Yuan Wu, Yi Chang
arXiv:2609.13228v1 Announce Type: new
Abstract: Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasonin...
By Marko Jojic, Zhaonan Li, Ben Zhou
arXiv:2606. 05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) still fail at causal physical reasoning, often producing plausible but incorrect answers.
By Tianyi Tang, Zhuoyi Lin, Zeyu Feng, Tianyi Ma, Yew-Soon Ong, Ivor Tsang, Haiyan Yin
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
By Peng, Lee, Yin Zhang, Yanglin Zhang, Haonan Wu, Zishan Liu, Ruoxi Zang, Xin Zhu, Jiayin Zheng, Jian Yao, Zefeng Ji, Fei Ma
arXiv:2609.23184v1 Announce Type: new
Abstract: Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly en...
By Ziming Xu, Shuang Liang, Ruobing Han, Ziqiao Xi, Mingxing Rao, Kun Zhou, Zijun Zhang, Yuchen Yan, Yufan Wei, Junbo Huang, Yifei Shao, Fang Nan, Biwei Huang
arXiv:2608.22429v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this relian...
By Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie