arXiv Computer Vision

PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image

PhysVGGT is a feed‑forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, along with object‑level mass, from a single RGB image in one forward pass. It treats physical property estimation as a dense per‑pixel prediction problem, using a visual geometry transformer to extract geometry‑aware tokens and separate dense and global prediction branches. A scalable pseudo‑label generation pipeline enables large‑scale weakly supervised training, and the model achieves state‑of‑the‑art performance on the ABO‑500 dataset while running 27× faster than previous methods.

Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.

arXiv Computer Vision
Sep 17

StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions

arXiv:2609.18430v1 Announce Type: new Abstract: Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucP...

By WM Team, Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang, Kaicheng Yu
arXiv AI
Jul 10

Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments

arXiv:2607. 07885v1 Announce Type: cross Abstract: Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical.

By Erik Jagnandan, Mulugeta Haile, Gregory Barber, Pratik Chaudhari
arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang