Feed-forward 3D foundation models such as VGGT predict cameras, depth, and point maps in a single pass, but can fail silently under low overlap, low parallax, and extreme relative rotation. Stratified...
The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.
By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua
GeoCond is a lightweight reliability adapter that enhances frozen feed‑forward 3D reconstruction backbones by reading their predicted geometry to produce pose‑level uncertainty and a refinement gate. It can be trained using permutation‑orbit variance, ground‑truth pose error, or cycle residuals from unlabeled pose graphs, and at inference requires only a single backbone pass plus a small MLP. On the VGGT backbone, GeoCond reduces out‑of‑distribution AUSE from 0.32 to 0.20, transfers zero‑shot to outdoor extreme‑view scenes, and prevents collapse from uniform bundle adjustment, while also enabling gated refinement, pose‑graph weighting, calibration, curation, and capture decisions.
By David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida, Lars Petersson
arXiv:2605.13018v2 Announce Type: replace
Abstract: Object-centric scene understanding is a fundamental challenge in computer vision. Existing approaches often rely on multi-stage pipelines that firs...
By Yi Du, Yang You, Xiang Wan, Leonidas Guibas
arXiv:2609.15795v2 Announce Type: replace
Abstract: Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental...
By Mingkai Liu, Hao Zhao, Xingxing Zuo
arXiv:2609.15795v1 Announce Type: new
Abstract: Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental issu...
By Mingkai Liu, Hao Zhao, Xingxing Zuo
arXiv:2508.03077v2 Announce Type: replace
Abstract: Feedforward 3D Gaussian Splatting (3DGS) overcomes the limitations of optimization-based 3DGS by enabling fast and high-quality reconstruction with...
By Anran Wu, Long Peng, Xin Di, Xueyuan Dai, Chen Wu, Yang Wang, Xueyang Fu, Yang Cao, Zheng-Jun Zha
arXiv:2608. 07579v1 Announce Type: cross Abstract: The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only.
By Abdullah Naeem, Anav Katwal, Ayon Dey, Noman Khan, Md Tamjidul Hoque
arXiv:2609.23733v1 Announce Type: new
Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
By Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai
arXiv:2609.14634v1 Announce Type: new
Abstract: Recently, 3D Gaussian Splatting SLAM (3DGS-SLAM) has gained significant momentum in simultaneous localization and 3DGS scene reconstruction. In real-wo...
By Kumaran Karthik, Pramat Shastri Jois, Suresh Sundaram
Matisse is a training‑free framework that combines active 3D reconstruction with keyframe selection by using evidence from a pretrained generative 3D model. It estimates evidential uncertainty via cross‑attention on 3D latent tokens and derives an evidential information gain to guide view acquisition and keyframe selection, reducing redundant observations and supporting multi‑object scenes with occlusion‑aware aggregation. On GSO30, YCB‑V, and Replica, Matisse improves Chamfer distance by 12.7%, 3.8%, and 9.2% respectively, and speeds up end‑to‑end reconstruction by 1.5× compared to the best baseline.
By Xihang Yu, Kaichen Zhou, Lorenzo Shaikewitz, Cl\'ement Jambon, Xiao Zhan, Rajat Talak, Luca Carlone
The paper argues that internal self-consistency checks cannot guarantee the accuracy of photogrammetric reconstructions, a limitation that is structural rather than a tuning issue. It introduces a track‑leakage‑free hold‑out protocol that withholds a deterministic subset of images and tests each against only 3D points supported by at least two retained images, ensuring no view is evaluated against the structure it helped create. Experiments on diverse datasets show that while the protocol is well‑posed, it saturates at a confidence score of 1.00 and fails to detect coherent distortion, missing large errors that can reach over 100 m.
whyItMatters":"The study highlights that hold‑out self‑validation scores, increasingly used as quality evidence for metric deliverables, may be misleading and cannot replace external survey validation."
By Behnam Asadi