Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Predicting object dynamics (i. e.
PhysVGGT is a feed‑forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, along with object‑level mass, from a single RGB image in one forward pass. It treats physical property estimation as a dense per‑pixel prediction problem, using a visual geometry transformer to extract geometry‑aware tokens and separate dense and global prediction branches. A scalable pseudo‑label generation pipeline enables large‑scale weakly supervised training, and the model achieves state‑of‑the‑art performance on the ABO‑500 dataset while running 27× faster than previous methods.
Predicting object dynamics (i. e.
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.
arXiv:2512. 10668v2 Announce Type: replace Abstract: A deep understanding of the physical world is essential for robotic manipulation and physically realistic simulation.
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
arXiv:2609.38620v1 Announce Type: new Abstract: Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-...
arXiv:2608. 20009v1 Announce Type: new Abstract: Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion.
arXiv:2609.18430v1 Announce Type: new Abstract: Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucP...
arXiv:2609.19142v1 Announce Type: new Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse...
arXiv:2606. 26700v1 Announce Type: cross Abstract: Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation.
arXiv:2607. 07885v1 Announce Type: cross Abstract: Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical.
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into do...