arXiv:2610.07569v1 Announce Type: cross
Abstract: Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping...
By Binh Long Nguyen, Kien Nguyen, Clinton Fookes, Peyman Moghadam
arXiv:2606. 06721v1 Announce Type: cross Abstract: Robots that operate over extended periods should not merely visit space; they should progressively understand it.
By Junyu Mao, Sara Ayoubi, Vishnu D. Sharma, Ilija Had\v{z}i\'c, Matthew Andrews
arXiv:2601. 10168v3 Announce Type: replace-cross Abstract: Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image aggregation under occlusions and constrained viewpoints.
By Yue Chang, Rufeng Chen, Zhaofan Zhang, Yi Chen, Yifan Tian, Sihong Xie
The paper introduces SceneLM, a vision‑language model that maintains an open‑vocabulary 3D scene map using only a structured text list of objects as persistent memory. The model updates this textual map by adding, editing, and removing objects for each input image, learning the process through supervision tasks and an automatic annotation pipeline. Evaluations on language‑grounded retrieval and localization benchmarks show competitive performance with traditional mapping systems while producing a 6‑12× more compact representation, and the model can run online on an edge device such as a quadruped robot.
By Adam Lilja, Fabio H\"ubel, Siming He, Junsheng Fu, Claire Tomlin, Lars Hammarstrand, Jitendra Malik, Jonas Frey, Marco Pavone
arXiv:2605.25059v4 Announce Type: replace
Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Em...
By Ruoyu Wang, Yong Liu, Jiahan Li, Sheng Tao, Yuhang Lin, Yukai Ma
arXiv:2512. 05131v2 Announce Type: replace-cross Abstract: Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images.
By Tianling Xu, Shengzhe Gan, Leslie Gu, Yuelei Li, Fangneng Zhan, Hanspeter Pfister
arXiv:2609.38620v1 Announce Type: new
Abstract: Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-...
By Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu, Mathias Otnes, Jeffrey Delmerico, Rui Wang, Yulun Tian, Nikolay Atanasov
Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments.
LangMap introduces a human‑verified benchmark for language‑conditioned goal navigation (LGN) that spans four hierarchical semantic levels—scene, room, region, and instance—within real‑world indoor 3D scans. The dataset, built on HM3D, contains 18K tasks with concise and detailed descriptions for 414 object categories, and its contrastive annotation protocol ensures high‑quality, discriminative region and instance labels. Evaluation shows that LangMap’s descriptions improve text‑to‑view matching accuracy by 23 points over GOAT‑Bench and achieve a 92.5% unique‑and‑correct match rate in an independent human audit, while a proposed RGB‑only baseline, PlaNaVid, attains top‑tier success rates without depth or 3D scene representations.
By Bo Miao, Weijia Liu, Jun Luo, Lachlan Shinnick, Jian Liu, Thomas Hamilton-Smith, Yuhe Yang, Zijie Wu, Vanja Videnovic, Feras Dayoub, Anton van den Hengel
Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.
arXiv:2607. 11498v1 Announce Type: cross Abstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions.
By Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Minho Park, Jaegul Choo
arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.
By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu