arXiv:2608.31025v1 Announce Type: new
Abstract: Inferring object dynamics from visual observations is essential for intelligent agents to reason about and interact with the physical world, yet remain...
By Jailing Lin, Jikuan Zhang, Jianhua Sun
SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph to provide geometric grounding and inter‑object relationships for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over existing methods, enabling physically interactive mixed reality experiences without manual tuning.
By Suji Kang, Seok-Young Kim, Young Bin Kim, Taewook Ha, Dieter Schmalstieg, Shohei Mori, Woontack Woo
SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—such as mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph that encodes inter‑object relationships, providing structured context for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over prior methods, reducing errors in mass estimation and enhancing scene‑level F‑Score.
arXiv:2609.13152v1 Announce Type: new
Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experim...
By Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu, Robin Jeshua Deepak, Stefano Saravalle, Aditya Shah
OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.
By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.