How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position
arXiv:2606. 15032v2 Announce Type: replace Abstract: World models have become a central abstraction in modern AI.
arXiv:2606. 15032v1 Announce Type: new Abstract: World models have rapidly become one of the central abstractions in modern AI.
arXiv:2606. 15032v2 Announce Type: replace Abstract: World models have become a central abstraction in modern AI.
arXiv:2609.05834v1 Announce Type: new Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the...
The paper surveys 160 benchmarks from 2017‑2026 that evaluate predictive embodied intelligence, categorising them into policy suites, embodied agents, world‑model evaluation, and prediction‑to‑action bridges. It finds that most benchmarks are model‑agnostic, rarely compare Vision‑Language‑Action policies to world models, and seldom turn predictions into executed actions. The authors argue that the lack of benchmarks designed to directly test the closed‑loop advantage of world models prevents the field from answering whether such models truly improve robotic performance.
arXiv:2608.24885v1 Announce Type: cross Abstract: Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on a...
arXiv:2511.09057v4 Announce Type: replace-cross Abstract: A world model is a cognitive simulator of the real-world environment allowing biological agents to reason about how the world evolves, whethe...
arXiv:2609.39235v1 Announce Type: cross Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
arXiv:2608. 09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation.
WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.
PAVXploreRL introduces a reinforcement learning framework that builds on a pretrained latent world model to explicitly optimize Physical Plausibility, Action Adherence, and Visual Fidelity (PAV) objectives. By combining in‑distribution expert trajectories with noise‑driven out‑of‑distribution action exploration, the method avoids reliance on paired video supervision and improves generalization. Experiments demonstrate a 5.6% average performance gain over pretrained baselines and more reliable policy evaluation with reduced overestimation bias.
arXiv:2607. 04681v1 Announce Type: cross Abstract: Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models.
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithful...
arXiv:2607. 00836v1 Announce Type: cross Abstract: World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities.