Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.
arXiv:2608.22102v1 Announce Type: cross
Abstract: We present GCA (Gaussian Constitutive Alignment), a framework for learning implicit constitutive laws from monocular dynamic video of deformable obje...
By Xiaoyang Liu, Kai Han
arXiv:2503. 24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation.
By Mikel Zhobro, Andreas Ren\'e Geist, Georg Martius
arXiv:2603. 03485v3 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models.
By Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang, Maojiang Su, Guo Ye, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Zhaoran Wang, Han Liu
arXiv:2606. 15015v1 Announce Type: cross Abstract: Physics-grounded video generation requires controllable 3D object dynamics that remain physically consistent under contact, deformation, and external forcing.
By Qizhen Ying, Guangming Wang, Yangchen Pan, Victor Adrian Prisacariu, Yixiong Jing
Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding.
arXiv:2609.01059v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. How...
By Jiayu Ding, Zhuodong Liu, Lei Zhang, Manyu Xiong, Hongbo Jin, Haoran Tang, Hongbo Zhang, Changen Zhu, Wenbo Xing
GeoWAM introduces a visual geometry world action model that predicts future scene geometry instead of future images, using point clouds to capture spatial structure and transformations. The model is pretrained to forecast geometry, then a geometry-conditioned action head predicts ego trajectories. Experiments show that this geometry-based approach yields stronger driving policies than image-based alternatives.
By Yiren Lu, Xin Ye, Jiaming Liu, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman
arXiv:2506. 06006v3 Announce Type: replace-cross Abstract: Can unified vision-language models (VLMs) perform forward dynamics prediction (FDP), i.
By Yifu Qiu, Yftah Ziser, Anna Korhonen, Shay B. Cohen, Edoardo M. Ponti
arXiv:2602. 06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex sensory information, such as visual perception, in situations where standard simulators struggle.
By Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, Ludovic Righetti
Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assumption that often fails for real-world in-the-wild videos.
arXiv:2609.01551v1 Announce Type: new
Abstract: Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations en...
By Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan, Sonia Joseph, Matthew Kowal, Konstantinos G. Derpanis