Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
arXiv:2603. 03485v3 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models.
arXiv:2606. 01538v1 Announce Type: cross Abstract: To study the ability to infer physical dynamics from videos and extrapolate them forward in time, we assemble a dataset of 2D Material Point Method (MPM) physical simulations covering rich physical phenomena such as deformable objects, fluids, kinetic objects, and emitters.
arXiv:2603. 03485v3 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models.
arXiv:2606. 05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators.
PhysPlan is a training‑free guidance framework that enhances video diffusion models by incorporating physical awareness through agentic physics simulation. It uses a vision‑language model to generate a Chain‑of‑Visual‑Thought representation of kinematic trajectories and 3D depth, which then drives an object‑centric test‑time optimization that isolates kinematic changes and locks the passive environment. The framework also employs Kinetic Intensity Profiling to adapt hyperparameters to varying physical deformations, and demonstrates superior performance on PhyGenBench and Physics‑IQ benchmarks compared to existing VDM baselines.
arXiv:2608.31025v1 Announce Type: new Abstract: Inferring object dynamics from visual observations is essential for intelligent agents to reason about and interact with the physical world, yet remain...
The paper introduces CoDeR, a new paradigm for world modeling that explicitly builds an executable world using code rather than relying solely on visual observations. CoDeR translates high‑level concepts into structured world rules, executable dynamics, and perceptual observations through five complementary roles, enabling long‑term memory, open‑ended interactions, autonomous world evolution, and persistent multi‑agent dynamics. Experiments show that this framework extends the capabilities of existing world models and achieves state‑of‑the‑art performance across multiple evaluation settings.
arXiv:2508.13009v5 Announce Type: replace Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...
arXiv:2607. 25321v1 Announce Type: new Abstract: Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity.
The paper introduces a physics‑informed diffusion guidance technique that uses self‑generated data augmentation to condition the diffusion model on the deviation from physical laws. By setting this deviation to zero during sampling, the method decouples equation evaluation from training and sampling, eliminating the need to solve governing equations at each denoising step. Experiments show the approach substantially reduces deviations from true dynamics and further improves performance when combined with existing physics‑constrained diffusion methods.
CompAdapt is a physics-consistent text-to-video generation framework that extends diffusion-based models to handle composite physical behaviors such as coupled motions, multi-stage transitions, and multi-object collisions. It translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial parameters. The system introduces dynamics-aware prior matching for one-shot adaptation to new physical environments and a physics-aware latent feature fusion module to enhance visual fidelity during fast, complex motion, outperforming existing physics-constrained baselines on physics-focused T2V benchmarks.
arXiv:2606. 00115v1 Announce Type: cross Abstract: Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models.
FracGen is a fracture‑aware video generation model that creates realistic, controllable fracture dynamics from a single intact image, guided by physics signals. It is trained using FracSim, a simulation framework that extends material point method (MPM) with a continuum damage model to produce paired fracture videos and dense physical fields. The model jointly predicts RGB video and physical maps, employing physics‑informed losses to capture material‑specific fracture behavior and enabling fine‑grained control over tear location, crack speed, and deformation before failure.
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.