arXiv AI

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

arXiv:2512. 05131v2 Announce Type: replace-cross Abstract: Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images.

arXiv Computer Vision
3d ago

Matisse: Evidence-Space Reasoning for Active 3D Reconstruction

Matisse is a training‑free framework that combines active 3D reconstruction with keyframe selection by using evidence from a pretrained generative 3D model. It estimates evidential uncertainty via cross‑attention on 3D latent tokens and derives an evidential information gain to guide view acquisition and keyframe selection, reducing redundant observations and supporting multi‑object scenes with occlusion‑aware aggregation. On GSO30, YCB‑V, and Replica, Matisse improves Chamfer distance by 12.7%, 3.8%, and 9.2% respectively, and speeds up end‑to‑end reconstruction by 1.5× compared to the best baseline.

By Xihang Yu, Kaichen Zhou, Lorenzo Shaikewitz, Cl\'ement Jambon, Xiao Zhan, Rajat Talak, Luca Carlone
arXiv Computer Vision
Aug 28

NeuDonatello: Uncertainty-Aware Framework for Accurate Neural SDF Learning

NeuDonatello is a new framework for neural signed distance function (SDF) learning that explicitly models spatially varying uncertainty using Monte Carlo sampling. By incorporating this uncertainty into an adaptive regularization scheme and an uncertainty-aware SDF-to-density conversion, the method selectively strengthens geometric constraints where RGB supervision is unreliable, thereby improving surface reconstruction accuracy. Experiments show that NeuDonatello achieves state‑of‑the‑art results on diverse scenes using only posed RGB images.

By Alvin Jinsung Choi, Wanhee Kim, Taeyun Kim, Dasol Hong, Wooju Lee, Hyun Myung
arXiv AI
Aug 11

RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation

arXiv:2601. 10168v3 Announce Type: replace-cross Abstract: Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image aggregation under occlusions and constrained viewpoints.

By Yue Chang, Rufeng Chen, Zhaofan Zhang, Yi Chen, Yifan Tian, Sihong Xie
arXiv AI
Jul 1

MVP-Nav: Multi-layer Value Map Planner Navigator

arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.

By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu
arXiv Computer Vision
Sep 22

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.

By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua