OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments
arXiv:2606. 29786v2 Announce Type: replace Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments.
arXiv:2601. 10168v3 Announce Type: replace-cross Abstract: Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image aggregation under occlusions and constrained viewpoints.
arXiv:2606. 29786v2 Announce Type: replace Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments.
arXiv:2606. 06721v1 Announce Type: cross Abstract: Robots that operate over extended periods should not merely visit space; they should progressively understand it.
arXiv:2607. 00889v1 Announce Type: cross Abstract: We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences.
arXiv:2609.22351v1 Announce Type: new Abstract: Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a ro...
arXiv:2610.01744v1 Announce Type: new Abstract: Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented ro...
Indoor visual relocalization plays a critical role in emerging spatial and embodied AI applications. However, prior research was predominantly devoted to low-level vision schemes, struggling to perceive scene semantics and compositions, which limits both interpretability and applicability.
arXiv:2608.21136v1 Announce Type: new Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...
arXiv:2610.01758v1 Announce Type: new Abstract: Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene...
Scene-Q is a confidence‑aware, coarse‑to‑fine querying framework for indoor mobile robots that grounds natural‑language queries in a 3D map. It normalizes encoder scores with temperature scaling and only invokes a reasoning VLM for low‑confidence cases, while high‑confidence queries are answered by fast retrieval. The method improves open‑vocabulary 3D instance segmentation on ScanNet200 and natural‑language 3D instance retrieval on real‑world reconstructions, especially for spatial and relational queries, while maintaining a substantial fraction of queries on the fast path.
The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.
The paper introduces 3DWay, a method that predicts 3D-consistent waypoints for robot manipulation by first generating multi‑view consistent 2D waypoints and then triangulating them. This approach addresses the 3D ambiguity inherent in 2D trajectory predictions and leverages pretrained vision‑language models to provide explicit 3D motion specifications. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing generalization for robot manipulation tasks.
Metric-Bench introduces a new benchmark for Vision‑Language Models (VLMs) that focuses on metric‑spatial reasoning in indoor scenes by using in‑image reference objects with known dimensions. The accompanying MetricReasoner fine‑tuning recipe employs structured prompts and numerical rewards to implicitly learn 2D‑to‑3D mapping without camera intrinsics. Experiments show that this approach improves spatial metric understanding by 43.1 % over existing models and boosts downstream embodied tasks, while also delivering gains on general VLM benchmarks.