arXiv AI
Sep 25

Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners

The paper argues that current embodied vision‑language planning benchmarks favor linguistic next‑token prediction over physically grounded next‑state reasoning, leading models to rely on language priors rather than true causal dependencies. To address this, the authors introduce Causal‑Plan‑Bench, a diagnostic suite covering four causal dimensions, and Causal‑Plan‑1M, a million‑scale corpus of explicit causal reasoning traces extracted from egocentric videos. Extensive experiments show that existing models perform poorly on these tasks, while a new model trained with a tailored recipe—Causal Planner based on Qwen3‑VL‑8B—achieves significant gains, demonstrating the feasibility of physically grounded causal reasoning.

By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Zirui Song, Lijie Wang, Chong Luo, Bei Liu, Yiming Li