arXiv:2609.01530v1 Announce Type: new
Abstract: Self-supervised pre-training via cross-view completion learns strong features for 3D vision from co-visible regions of image pairs. However, the refere...
By Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit
arXiv:2606. 09919v1 Announce Type: cross Abstract: Perceptual uncertainty is a central challenge for heterogeneous robot teams operating in unstructured outdoor environments, where no single viewpoint affords reliable scene understanding.
By Michal P. Podolinsky, Neel P. Bhatt, Pranay Samineni, Rohan Siva, Christian Ellis, Ufuk Topcu
arXiv:2512. 23020v3 Announce Type: replace-cross Abstract: 3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes.
By Wenyuan Huang, Zhenyu Zhang, Zhao Wang, Zhou Wei, Ting Huang, Fang Zhao, Jian Yang
arXiv:2607. 16012v1 Announce Type: cross Abstract: Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation.
By Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim
AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.
By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv:2609.23733v1 Announce Type: new
Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
By Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai
IVSGround introduces a lightweight view selector that learns to choose the most informative camera views for vision‑language model (VLM) based 3D visual grounding, replacing heuristic view selection. The selector is trained via a two‑stage rejection sampling process that uses feedback from a reasoning VLM to generate supervision signals. Experiments on ScanRefer and NR3D demonstrate that IVSGround consistently improves grounding accuracy over existing zero‑shot pipelines, underscoring the importance of selecting where to look for effective 3D visual grounding.
By Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu, Chun-Wei Huang, Quan Kong, Chun-Yi Lee
SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.
By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem
As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination.
arXiv:2605.25059v4 Announce Type: replace
Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Em...
By Ruoyu Wang, Yong Liu, Jiahan Li, Sheng Tao, Yuhang Lin, Yukai Ma
arXiv:2606.03406v2 Announce Type: replace
Abstract: Reliable correspondence estimation supports image processing and 3D vision tasks, including Structure from Motion, visual localization, and image r...
By Xu Pan, Zhen Pang, Qiyuan Ma, Wei Ji, Shuhan Shen, Xianwei Zheng
arXiv:2610.03015v1 Announce Type: cross
Abstract: Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors....
By Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang