HappyWorld-Bench
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWor...
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWor...
RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.
RobotEQ-Video is a new video-centric benchmark designed to advance Social Proactive Intelligence (SPI) by moving beyond static image analysis. It introduces a hierarchical world-state taxonomy with 6 domains, 20 dimensions, 142 level‑1 attributes, and 816 level‑2 attributes, and includes over 2,000 videos annotated with 100,000+ human labels and 16,000+ behavior‑properness tags. Evaluation shows existing systems underperform humans, highlighting the need for richer video data and comprehensive scenario coverage in SPI research.
arXiv:2608. 09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation.
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs.
arXiv:2511. 17649v4 Announce Type: replace-cross Abstract: Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday human-built environments.
arXiv:2608.16859v2 Announce Type: replace Abstract: A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especi...
arXiv:2606. 15032v2 Announce Type: replace Abstract: World models have become a central abstraction in modern AI.
WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2606. 10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood.
arXiv:2606. 15032v1 Announce Type: new Abstract: World models have rapidly become one of the central abstractions in modern AI.