arXiv AI

From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal Reasoning

arXiv:2606. 11745v1 Announce Type: cross Abstract: Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual inputs and reasoning over intervention effects.

arXiv Computer Vision
6d ago

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

CCRV-Bench is a constraint‑driven benchmark designed to evaluate visual causal reasoning in vision‑language models on single‑image physical scenarios. It assesses four causal task dimensions—causal relation discovery, state prediction, causal diagnosis, and intervention—while applying constraints such as entity symbolization, spatial grounding, factual adversarial constraints, and minimalist output constraints to reduce shortcut learning. Experiments on 15 multimodal models reveal that constraint sensitivity varies by task and model, with intervention and spatial grounding having the largest impact and factual adversarial constraints improving causal diagnosis across models.

By Linyuan Gao, Yuan Wu, Yi Chang
arXiv AI
Jun 6

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

arXiv:2606. 05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) still fail at causal physical reasoning, often producing plausible but incorrect answers.

By Tianyi Tang, Zhuoyi Lin, Zeyu Feng, Tianyi Ma, Yew-Soon Ong, Ivor Tsang, Haiyan Yin
arXiv Computer Vision
Sep 22

CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model

arXiv:2609.23184v1 Announce Type: new Abstract: Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly en...

By Ziming Xu, Shuang Liang, Ruobing Han, Ziqiao Xi, Mingxing Rao, Kun Zhou, Zijun Zhang, Yuchen Yan, Yufan Wei, Junbo Huang, Yifei Shao, Fang Nan, Biwei Huang
arXiv AI
Sep 25

Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners

The paper argues that current embodied vision‑language planning benchmarks favor linguistic next‑token prediction over physically grounded next‑state reasoning, leading models to rely on language priors rather than true causal dependencies. To address this, the authors introduce Causal‑Plan‑Bench, a diagnostic suite covering four causal dimensions, and Causal‑Plan‑1M, a million‑scale corpus of explicit causal reasoning traces extracted from egocentric videos. Extensive experiments show that existing models perform poorly on these tasks, while a new model trained with a tailored recipe—Causal Planner based on Qwen3‑VL‑8B—achieves significant gains, demonstrating the feasibility of physically grounded causal reasoning.

By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Zirui Song, Lijie Wang, Chong Luo, Bei Liu, Yiming Li
arXiv Computer Vision
Aug 25

Investigating Relational Reasoning in VLMs

arXiv:2608.23518v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or...

By Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap
arXiv Computation and Language
Sep 7

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

The study examines how Vision‑Language Models (VLMs) integrate visual evidence into language‑based decisions by applying layer‑wise causal interventions on video‑text attention pathways in a video‑based generative multiple‑choice setting. Findings reveal that visual information is primarily incorporated while processing candidate answer options, with nouns serving as key semantic anchors and verbs becoming important during temporal reasoning. The research also uncovers a distinct pattern in temporal reasoning, indicating that VLMs struggle to reconstruct sequential information across video frames, possibly due to linguistic biases in temporal expressions.

By Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt
arXiv AI
Aug 10

Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning

arXiv:2608. 06938v1 Announce Type: cross Abstract: The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions.

By Chen Ling, Hanqian Li, Dongnan Liu, Keyu Qian, Jungang Li, Xinglong liu, Shiyi Wang, Xin Dong, Pengcheng Zhu, Wei Zhou, Linjian Mo, Nai Ding