arXiv AI

From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs

arXiv:2606. 09134v1 Announce Type: cross Abstract: Constructing knowledge graphs from 3D simulation scenes is essential for robot task reasoning, but the key bottleneck, grounding scene objects to formal ontology classes, still relies on manually curated dictionaries that are brittle and do not generalize across assets.

arXiv AI
Jul 21

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering

arXiv:2510. 01483v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long egocentric video remains costly: because VLMs maintain no persistent memory or explicit spatial representation, all sampled frames must be re-processed for every new query.

By Mohamad Al Mdfaa, Svetlana Lukina, Timur Akhtyamov, Arthur Nigmatzyanov, Dmitrii Nalberskii, Sergey Zagoruyko, Gonzalo Ferrer
arXiv AI
Sep 10

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

GoDeep is an annotation‑free method for open‑vocabulary 3D scene understanding that uses a vision‑language model solely as a translator to generate structured, entity‑level descriptions of each image. These descriptions are projected and aggregated in a language‑only embedding space, eliminating the need for a 3D training corpus or domain‑specific encoder. The approach achieves competitive performance on ScanNet++ and a cultural heritage benchmark, accurately localizes out‑of‑vocabulary objects, and offers explainable, point‑level predictions.

By Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
Hugging Face Trending Papers
Aug 6

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments.

arXiv Computer Vision
Sep 18

CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

CitySTAR introduces a training‑free framework that transforms billion‑scale urban point clouds into a query‑ready scene graph of open‑vocabulary 3D instances, using CodeLLM‑driven tools to supply multimodal evidence for node attributes and spatial relations. It models target‑context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation, followed by a Reflective Cross‑modal Grounding module that integrates topology consistency and 2D visual evidence to decide over a metric‑aware 3D context graph. The authors also present CitySTAR‑3D, a benchmark that enhances semantic coverage, instance completeness, bounding‑box fidelity, and spatial‑relation complexity for city‑scale 3D grounding, and report extensive experiments showing consistent improvements in open‑world urban 3D grounding with strong interpretability and generalization.

By Shuai Zhang, Hongye Hou, Qinghe Liu, Zhuoxiao Li, Dongli Wu, Jing Ou, Yuan Liu, Wufan Zhao