arXiv Machine Learning

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics

arXiv:2606. 01538v1 Announce Type: cross Abstract: To study the ability to infer physical dynamics from videos and extrapolate them forward in time, we assemble a dataset of 2D Material Point Method (MPM) physical simulations covering rich physical phenomena such as deformable objects, fluids, kinetic objects, and emitters.

arXiv Computer Vision
Sep 14

Physics-Aware Video Generation via Agentic Planning and Graph-Guided Optimization

PhysPlan is a training‑free guidance framework that enhances video diffusion models by incorporating physical awareness through agentic physics simulation. It uses a vision‑language model to generate a Chain‑of‑Visual‑Thought representation of kinematic trajectories and 3D depth, which then drives an object‑centric test‑time optimization that isolates kinematic changes and locks the passive environment. The framework also employs Kinetic Intensity Profiling to adapt hyperparameters to varying physical deformations, and demonstrates superior performance on PhyGenBench and Physics‑IQ benchmarks compared to existing VDM baselines.

By Minh-Loi Nguyen, Xuan-Vu Le, Thanh-Toan Do, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
arXiv Computer Vision
Sep 23

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

The paper introduces CoDeR, a new paradigm for world modeling that explicitly builds an executable world using code rather than relying solely on visual observations. CoDeR translates high‑level concepts into structured world rules, executable dynamics, and perceptual observations through five complementary roles, enabling long‑term memory, open‑ended interactions, autonomous world evolution, and persistent multi‑agent dynamics. Experiments show that this framework extends the capabilities of existing world models and achieves state‑of‑the‑art performance across multiple evaluation settings.

By Zixun Fang, Yawen Shao, Kai Zhu, Jie Xiao, Shihan Chen, Yu Liu, Xueyang Fu, Yang Cao, Wei Zhai, Zheng-Jun Zha
arXiv Computer Vision
4d ago

Matrix-game 2.0: An open-source, real-time, and streaming interactive world model

arXiv:2508.13009v5 Announce Type: replace Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...

By Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou
arXiv AI
Jul 29

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

arXiv:2607. 25321v1 Announce Type: new Abstract: Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity.

By Ruijie Su, Yuanzhi Liang, Xiaohua Xie, Jianhuang Lai
Hugging Face Trending Papers
Aug 27

Self-Augmented Diffusion Guidance for Physics-Informed Generation

The paper introduces a physics‑informed diffusion guidance technique that uses self‑generated data augmentation to condition the diffusion model on the deviation from physical laws. By setting this deviation to zero during sampling, the method decouples equation evaluation from training and sampling, eliminating the need to solve governing equations at each denoising step. Experiments show the approach substantially reduces deviations from true dynamics and further improves performance when combined with existing physics‑constrained diffusion methods.

arXiv Computer Vision
Sep 21

CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation

CompAdapt is a physics-consistent text-to-video generation framework that extends diffusion-based models to handle composite physical behaviors such as coupled motions, multi-stage transitions, and multi-object collisions. It translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial parameters. The system introduces dynamics-aware prior matching for one-shot adaptation to new physical environments and a physics-aware latent feature fusion module to enhance visual fidelity during fast, complex motion, outperforming existing physics-constrained baselines on physics-focused T2V benchmarks.

By Haoran Qin (Harbin Institute of Technology, China), Renlong Wu (Harbin Institute of Technology, China), Tianyu Huang (Harbin Institute of Technology, China), Yukang Ding (Taobao, Alibaba Group, China), Hui Li (Harbin Institute of Technology, China), Wangmeng Zuo (Harbin Institute of Technology, China)
arXiv Computer Vision
4d ago

FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation

FracGen is a fracture‑aware video generation model that creates realistic, controllable fracture dynamics from a single intact image, guided by physics signals. It is trained using FracSim, a simulation framework that extends material point method (MPM) with a continuum damage model to produce paired fracture videos and dense physical fields. The model jointly predicts RGB video and physical maps, employing physics‑informed losses to capture material‑specific fracture behavior and enabling fine‑grained control over tear location, crack speed, and deformation before failure.

By Trong-Tung Nguyen, Jiahan Zhang, Anand Bhattad
Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.