OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments
arXiv:2606. 29786v2 Announce Type: replace Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments.
The paper introduces RelationVGGT, a feed‑forward framework that performs 3D spatial relation segmentation without per‑scene optimization or known camera poses. It combines semantic features from a visual foundation model with geometry‑aware representations from a 3D geometry foundation model, and uses a relation transformer to predict subject‑conditioned, cross‑view relations based on a visual subject and a textual query. The authors also present an automated annotation pipeline built on ScanNet++ with VLMs and LLMs to generate scalable training data for this new task.
arXiv:2606. 29786v2 Announce Type: replace Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments.
arXiv:2609.22687v1 Announce Type: new Abstract: We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg...
arXiv:2610.00040v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation mo...
Object-Uni is a unified model that integrates pose perception, spatial reasoning, pose-conditioned generation, and novel view synthesis for object-centric spatial understanding and controllable image generation. It treats object pose as an explicit geometric variable shared across tasks and introduces a viewpoint-based orientation abstraction to make pose interpretable by multimodal large language models. The authors also create a new benchmark, UniSpatial-80K, and demonstrate that Object‑Uni improves both pose understanding and pose‑controllable generation compared to existing models.
arXiv:2610.11810v1 Announce Type: new Abstract: Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to underst...
arXiv:2606. 27412v1 Announce Type: cross Abstract: 3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object-relation-object graphs, providing a compact relational abstraction for spatial understanding.
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding.
arXiv:2607. 00889v1 Announce Type: cross Abstract: We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences.
arXiv:2608. 15710v1 Announce Type: cross Abstract: We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison.
arXiv:2606. 03100v1 Announce Type: cross Abstract: Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial reasoning capabilities.
arXiv:2604. 02546v3 Announce Type: replace-cross Abstract: Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding.
arXiv:2601. 11729v2 Announce Type: replace-cross Abstract: Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems.