arXiv AI

RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation

arXiv:2601. 10168v3 Announce Type: replace-cross Abstract: Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image aggregation under occlusions and constrained viewpoints.

arXiv Computer Vision
Sep 18

Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning

Scene-Q is a confidence‑aware, coarse‑to‑fine querying framework for indoor mobile robots that grounds natural‑language queries in a 3D map. It normalizes encoder scores with temperature scaling and only invokes a reasoning VLM for low‑confidence cases, while high‑confidence queries are answered by fast retrieval. The method improves open‑vocabulary 3D instance segmentation on ScanNet200 and natural‑language 3D instance retrieval on real‑world reconstructions, especially for spatial and relational queries, while maintaining a substantial fraction of queries on the fast path.

By Juno Kim, Yesol Park, Hye-Jung Yoon, Byoung-Tak Zhang
Hugging Face Trending Papers
Sep 8

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.

arXiv AI
Sep 10

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

The paper introduces 3DWay, a method that predicts 3D-consistent waypoints for robot manipulation by first generating multi‑view consistent 2D waypoints and then triangulating them. This approach addresses the 3D ambiguity inherent in 2D trajectory predictions and leverages pretrained vision‑language models to provide explicit 3D motion specifications. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing generalization for robot manipulation tasks.

By Ziqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Yuxin Chen, Gu Wang, Xingyu Liu, Masayoshi Tomizuka, Xiangyang Ji
arXiv Computer Vision
Sep 23

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Metric-Bench introduces a new benchmark for Vision‑Language Models (VLMs) that focuses on metric‑spatial reasoning in indoor scenes by using in‑image reference objects with known dimensions. The accompanying MetricReasoner fine‑tuning recipe employs structured prompts and numerical rewards to implicitly learn 2D‑to‑3D mapping without camera intrinsics. Experiments show that this approach improves spatial metric understanding by 43.1 % over existing models and boosts downstream embodied tasks, while also delivering gains on general VLM benchmarks.

By Yuling Xi, Haokai Zhang, Muzhi Zhu, Hao Zhong, Zongze Du, Hengyu Zhao, Chenchen Jing, Yufei Yin, Bin Qin, Yongjie Yang, Zhenbo Luo, Hao Chen, Chunhua Shen