arXiv Computer Vision

Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos

Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.

arXiv Computer Vision
Aug 25

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM introduces a visual geometry world action model that predicts future scene geometry instead of future images, using point clouds to capture spatial structure and transformations. The model is pretrained to forecast geometry, then a geometry-conditioned action head predicts ego trajectories. Experiments show that this geometry-based approach yields stronger driving policies than image-based alternatives.

By Yiren Lu, Xin Ye, Jiaming Liu, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman
arXiv AI
Jun 3

Coupled Local and Global World Models for Efficient First Order RL

arXiv:2602. 06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex sensory information, such as visual perception, in situations where standard simulators struggle.

By Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, Ludovic Righetti