RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.
arXiv:2608.24885v1 Announce Type: cross Abstract: Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on a...
arXiv:2603. 22876v2 Announce Type: replace-cross Abstract: Learning a generalist control policy for robotic manipulation typically relies on large-scale datasets.
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithful...
arXiv:2607. 26903v1 Announce Type: new Abstract: The key bottleneck in embodied AI is not model architecture but data.
arXiv:2603. 10652v3 Announce Type: replace-cross Abstract: In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion.
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact.
arXiv:2606. 26904v1 Announce Type: cross Abstract: Video reasoning language models implicitly assume that every input frame is equally reliable.
arXiv:2603. 25937v2 Announce Type: replace-cross Abstract: Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations.