PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics
arXiv:2607. 20653v1 Announce Type: cross Abstract: Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge.
arXiv:2608. 20009v1 Announce Type: new Abstract: Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion.
arXiv:2607. 20653v1 Announce Type: cross Abstract: Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge.
JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
PhysVGGT is a feed‑forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, along with object‑level mass, from a single RGB image in one forward pass. It treats physical property estimation as a dense per‑pixel prediction problem, using a visual geometry transformer to extract geometry‑aware tokens and separate dense and global prediction branches. A scalable pseudo‑label generation pipeline enables large‑scale weakly supervised training, and the model achieves state‑of‑the‑art performance on the ABO‑500 dataset while running 27× faster than previous methods.
arXiv:2608.24044v1 Announce Type: new Abstract: Latent world models plan by predicting how candidate actions transform learned representations. In self-predictive models, however, the encoder and pre...
arXiv:2609.10464v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning,...
arXiv:2606. 28455v1 Announce Type: cross Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics.
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.
arXiv:2608.31025v1 Announce Type: new Abstract: Inferring object dynamics from visual observations is essential for intelligent agents to reason about and interact with the physical world, yet remain...
The paper introduces a sparse, residual world model that focuses on predicting only the changes in a scene by using a per-object change gate and a residual delta head. On a MuJoCo tabletop pushing benchmark, this approach outperforms a dense multilayer perceptron, achieving 2.5 to 4.6 times better next‑state pose accuracy with 8.6 to 11.1 times fewer parameters, maintaining high change‑detection F1 scores, and showing strong transfer across object counts. In autoregressive rollout and sampling‑based planning, the sparse model accumulates less error and enables successful planning where dense models fail.
arXiv:2608. 09876v1 Announce Type: cross Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics.
arXiv:2609.36521v1 Announce Type: new Abstract: Physical-field reconstruction and forecasting depend on both measurement density and spatial layout, yet evaluation under a single observation pattern...
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.