PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.
arXiv:2608. 02150v2 Announce Type: replace-cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities.
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors.
arXiv:2606. 18586v1 Announce Type: cross Abstract: Physical events are not understood by their names alone, but by the causal state changes that compose them.
CompAdapt is a physics-consistent text-to-video generation framework that extends diffusion-based models to handle composite physical behaviors such as coupled motions, multi-stage transitions, and multi-object collisions. It translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial parameters. The system introduces dynamics-aware prior matching for one-shot adaptation to new physical environments and a physics-aware latent feature fusion module to enhance visual fidelity during fast, complex motion, outperforming existing physics-constrained baselines on physics-focused T2V benchmarks.
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws.
The paper introduces Physically Plausible Video Generation (PPVG), a method that generates videos consistent with physical laws by treating physical evolution as a chain of causally connected events. It employs three modules: Physics-driven Event Chain Reasoning to decompose phenomena into scene-graph events, Transition-aware Routed Keyframe Conditioning to guide keyframe synthesis for smooth transitions, and Physics-injected Contrastive Semantic Guidance to steer generation toward plausible dynamics. Experiments on multiple physics benchmarks show improved physical plausibility compared to prior approaches.
arXiv:2609.13006v2 Announce Type: replace Abstract: Video diffusion models (VDMs) synthesize photorealistic content, yet they often fail to follow the course that a physical phenomenon should take wi...
arXiv:2607. 26452v1 Announce Type: new Abstract: World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure.
arXiv:2603. 03485v3 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models.
arXiv:2602. 10840v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving.
arXiv:2606. 02800v1 Announce Type: cross Abstract: We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture.