arXiv AI By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Zirui Song, Lijie Wang, Chong Luo, Bei Liu, Yiming Li

Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners

Read the original on arXiv AI →

The paper argues that current embodied vision‑language planning benchmarks favor linguistic next‑token prediction over physically grounded next‑state reasoning, leading models to rely on language priors rather than true causal dependencies. To address this, the authors introduce Causal‑Plan‑Bench, a diagnostic suite covering four causal dimensions, and Causal‑Plan‑1M, a million‑scale corpus of explicit causal reasoning traces extracted from egocentric videos. Extensive experiments show that existing models perform poorly on these tasks, while a new model trained with a tailored recipe—Causal Planner based on Qwen3‑VL‑8B—achieves significant gains, demonstrating the feasibility of physically grounded causal reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 6

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

arXiv:2606. 05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) still fail at causal physical reasoning, often producing plausible but incorrect answers.

By Tianyi Tang, Zhuoyi Lin, Zeyu Feng, Tianyi Ma, Yew-Soon Ong, Ivor Tsang, Haiyan Yin
arXiv Computer Vision
Sep 22

CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model

arXiv:2609.23184v1 Announce Type: new Abstract: Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly en...

By Ziming Xu, Shuang Liang, Ruobing Han, Ziqiao Xi, Mingxing Rao, Kun Zhou, Zijun Zhang, Yuchen Yan, Yufan Wei, Junbo Huang, Yifei Shao, Fang Nan, Biwei Huang
arXiv AI
Sep 24

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

arXiv:2609. 28366v1 Announce Type: cross Abstract: Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning.

By Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li
arXiv AI
Jul 7

Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

arXiv:2607. 04681v1 Announce Type: cross Abstract: Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models.

By Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone