CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 15869v1 Announce Type: cross Abstract: Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments.
The paper argues that current embodied vision‑language planning benchmarks favor linguistic next‑token prediction over physically grounded next‑state reasoning, leading models to rely on language priors rather than true causal dependencies. To address this, the authors introduce Causal‑Plan‑Bench, a diagnostic suite covering four causal dimensions, and Causal‑Plan‑1M, a million‑scale corpus of explicit causal reasoning traces extracted from egocentric videos. Extensive experiments show that existing models perform poorly on these tasks, while a new model trained with a tailored recipe—Causal Planner based on Qwen3‑VL‑8B—achieves significant gains, demonstrating the feasibility of physically grounded causal reasoning.
arXiv:2606. 01810v1 Announce Type: new Abstract: Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning.
arXiv:2603. 22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.
arXiv:2606. 11745v1 Announce Type: cross Abstract: Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual inputs and reasoning over intervention effects.
arXiv:2606. 15753v1 Announce Type: new Abstract: Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning.