HappyWorld-Bench
arXiv:2609.24308v1 Announce Type: new Abstract: Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, int...
arXiv:2609.24308v1 Announce Type: new Abstract: Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, int...
RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs.
arXiv:2608. 09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation.
RobotEQ-Video is a new video-centric benchmark designed to advance Social Proactive Intelligence (SPI) by moving beyond static image analysis. It introduces a hierarchical world-state taxonomy with 6 domains, 20 dimensions, 142 level‑1 attributes, and 816 level‑2 attributes, and includes over 2,000 videos annotated with 100,000+ human labels and 16,000+ behavior‑properness tags. Evaluation shows existing systems underperform humans, highlighting the need for richer video data and comprehensive scenario coverage in SPI research.
arXiv:2511. 17649v4 Announce Type: replace-cross Abstract: Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday human-built environments.
arXiv:2608.16859v2 Announce Type: replace Abstract: A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especi...
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.
arXiv:2606. 15032v2 Announce Type: replace Abstract: World models have become a central abstraction in modern AI.
arXiv:2606. 11909v1 Announce Type: new Abstract: Benchmarks are essential for evaluating embodied spatial intelligence, yet their construction is labor-intensive, hard to reuse, and difficult to maintain.
The paper surveys 160 benchmarks from 2017‑2026 that evaluate predictive embodied intelligence, categorising them into policy suites, embodied agents, world‑model evaluation, and prediction‑to‑action bridges. It finds that most benchmarks are model‑agnostic, rarely compare Vision‑Language‑Action policies to world models, and seldom turn predictions into executed actions. The authors argue that the lack of benchmarks designed to directly test the closed‑loop advantage of world models prevents the field from answering whether such models truly improve robotic performance.