World Models' Last Exam in Physics
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 05948v1 Announce Type: new Abstract: Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions.
Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated.
arXiv:2609.24308v1 Announce Type: new Abstract: Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, int...
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws.
arXiv:2608. 02150v2 Announce Type: replace-cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities.
PhysicsLENS is a new dataset and benchmark designed to evaluate how well video generation models capture physical properties relevant to robotics. It consists of matched scenario pairs that share the same conditioning frame and task but differ in underlying physics, and it covers seven physical domains such as collision, gravity, and friction. The benchmark includes over 400 human-annotated labels from four video generation models, revealing that many plausible-looking videos still ignore the specified physical property.