Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.
PhysPlan is a training‑free guidance framework that enhances video diffusion models by incorporating physical awareness through agentic physics simulation. It uses a vision‑language model to generate a Chain‑of‑Visual‑Thought representation of kinematic trajectories and 3D depth, which then drives an object‑centric test‑time optimization that isolates kinematic changes and locks the passive environment. The framework also employs Kinetic Intensity Profiling to adapt hyperparameters to varying physical deformations, and demonstrates superior performance on PhyGenBench and Physics‑IQ benchmarks compared to existing VDM baselines.
By Minh-Loi Nguyen, Xuan-Vu Le, Thanh-Toan Do, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
The paper introduces Physically Plausible Video Generation (PPVG), a method that generates videos consistent with physical laws by treating physical evolution as a chain of causally connected events. It employs three modules: Physics-driven Event Chain Reasoning to decompose phenomena into scene-graph events, Transition-aware Routed Keyframe Conditioning to guide keyframe synthesis for smooth transitions, and Physics-injected Contrastive Semantic Guidance to steer generation toward plausible dynamics. Experiments on multiple physics benchmarks show improved physical plausibility compared to prior approaches.
By Zixuan Wang, Yixin Hu, Wen Li, Feng Chen, Yan Liu, Duo Peng, Yinjie Lei
arXiv:2608. 04575v1 Announce Type: cross Abstract: Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions.
By Chen Yang, Shenxiang Zeng, Haoyang Zhao, Zhouyuan Xu, Youquan He, Haoyu Li, Mingyi Deng, Jiansheng Fan, Chen Wang
arXiv:2605.08712v2 Announce Type: replace
Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dime...
By Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, Qi Dou, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming Jin
arXiv:2607. 13451v1 Announce Type: cross Abstract: Simulating deformable objects is essential for a wide range of robotic manipulation applications, yet accurately predicting their dynamics remains challenging.
By Shivansh Patel, Kaifeng Zhang, Sanjay Pokkali, Svetlana Lazebnik, Yunzhu Li
arXiv:2607. 10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential.
By Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
arXiv:2607. 07601v1 Announce Type: cross Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations.
By Kaicong Huang, Meng Ma, Ruimin Ke
arXiv:2609.18430v1 Announce Type: new
Abstract: Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucP...
By WM Team, Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang, Kaicheng Yu
arXiv:2609.01059v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. How...
By Jiayu Ding, Zhuodong Liu, Lei Zhang, Manyu Xiong, Hongbo Jin, Haoran Tang, Hongbo Zhang, Changen Zhu, Wenbo Xing
arXiv:2610.02180v1 Announce Type: cross
Abstract: Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambig...
By Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad
arXiv:2610.00451v1 Announce Type: cross
Abstract: Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual...
By Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI)