Hugging Face Trending Papers

SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—such as mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph that encodes inter‑object relationships, providing structured context for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over prior methods, reducing errors in mass estimation and enhancing scene‑level F‑Score.

arXiv Computer Vision
Sep 18

SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

SnapPhysics is a training‑free framework that reconstructs 3D objects and estimates their physical properties—mass, friction, and center of gravity—from a single image. It combines instance‑level 3D reconstruction with a physics‑aware scene graph to provide geometric grounding and inter‑object relationships for vision‑language model reasoning. Experiments on 3D‑FRONT and real captured scenes show significant improvements over existing methods, enabling physically interactive mixed reality experiences without manual tuning.

By Suji Kang, Seok-Young Kim, Young Bin Kim, Taewook Ha, Dieter Schmalstieg, Shohei Mori, Woontack Woo
arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv AI
Jul 21

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering

arXiv:2510. 01483v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long egocentric video remains costly: because VLMs maintain no persistent memory or explicit spatial representation, all sampled frames must be re-processed for every new query.

By Mohamad Al Mdfaa, Svetlana Lukina, Timur Akhtyamov, Arthur Nigmatzyanov, Dmitrii Nalberskii, Sergey Zagoruyko, Gonzalo Ferrer
arXiv AI
Sep 4

GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

GraFT is a training‑free framework that enhances spatial reasoning in multimodal large language models by integrating a compact 3D scene graph (3DSG). It offers deterministic geometry via symbolic tools, allocentric layout through bird’s‑eye‑view rendering, and visual‑attribute grounding using egocentric frames. Experiments on ScanQA and VSI‑Bench show significant performance gains, with CIDEr increasing by 27% and improvements up to 65% over baseline models.

By Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu
arXiv AI
2d ago

PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

arXiv:2610.00451v1 Announce Type: cross Abstract: Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual...

By Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI)
arXiv AI
Aug 26

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

PhysMLLMs introduces physics-inspired spatial continuity priors into video multimodal large language models to address spatio‑temporal inconsistencies such as jitter, drift, and identity switches. The method, called Global Representation Prior Alignment (REPA‑Global), distills global visual representations from a frozen DINOv2 teacher during training, aligning student representations without affecting inference speed. Experiments on multiple video benchmarks show improved segmentation mask quality and cross‑frame consistency, especially for challenging scenarios involving small targets, fast motion, occlusion, and distractors, while maintaining comparable performance on single‑frame image segmentation and general VLM tasks.

By Siyao Yan, Bo Han, Jisheng Dang, Bimei Wang, Shude Wang, Hong Peng, Yulan Guo, Jianhuang Lai, Bin Hu, Tat-SengChua