Hugging Face Trending Papers

NewtPhys: Do Foundation Models Understand Newtonian Physics?

Read the original on Hugging Face Trending Papers →

Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 18

SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph to provide geometric grounding and inter‑object relationships for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over existing methods, enabling physically interactive mixed reality experiences without manual tuning.

By Suji Kang, Seok-Young Kim, Young Bin Kim, Taewook Ha, Dieter Schmalstieg, Shohei Mori, Woontack Woo
Hugging Face Trending Papers
Sep 17

SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—such as mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph that encodes inter‑object relationships, providing structured context for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over prior methods, reducing errors in mass estimation and enhancing scene‑level F‑Score.

arXiv Computation and Language
Aug 27

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.

By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.