arXiv AI

Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset

arXiv:2606. 16479v1 Announce Type: cross Abstract: Visual Geometry Grounded Transformer (VGGT) has already attracted a great deal of attention in a short period of time, not least due to the Best Paper Award at CVPR-2025.

arXiv Computer Vision
Sep 22

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.

By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua
arXiv Computer Vision
Sep 17

GeoCond: A Conditioning-Aware Reliability Adapter for Feed-Forward 3D Reconstruction

GeoCond is a lightweight reliability adapter that enhances frozen feed‑forward 3D reconstruction backbones by reading their predicted geometry to produce pose‑level uncertainty and a refinement gate. It can be trained using permutation‑orbit variance, ground‑truth pose error, or cycle residuals from unlabeled pose graphs, and at inference requires only a single backbone pass plus a small MLP. On the VGGT backbone, GeoCond reduces out‑of‑distribution AUSE from 0.32 to 0.20, transfers zero‑shot to outdoor extreme‑view scenes, and prevents collapse from uniform bundle adjustment, while also enabling gated refinement, pose‑graph weighting, calibration, curation, and capture decisions.

By David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida, Lars Petersson
arXiv Computer Vision
3d ago

Matisse: Evidence-Space Reasoning for Active 3D Reconstruction

Matisse is a training‑free framework that combines active 3D reconstruction with keyframe selection by using evidence from a pretrained generative 3D model. It estimates evidential uncertainty via cross‑attention on 3D latent tokens and derives an evidential information gain to guide view acquisition and keyframe selection, reducing redundant observations and supporting multi‑object scenes with occlusion‑aware aggregation. On GSO30, YCB‑V, and Replica, Matisse improves Chamfer distance by 12.7%, 3.8%, and 9.2% respectively, and speeds up end‑to‑end reconstruction by 1.5× compared to the best baseline.

By Xihang Yu, Kaichen Zhou, Lorenzo Shaikewitz, Cl\'ement Jambon, Xiao Zhan, Rajat Talak, Luca Carlone
arXiv Computer Vision
Sep 4

Hold-Out Self-Validation Cannot Certify Photogrammetric Accuracy: Saturation and Blindness to Coherent Distortion

The paper argues that internal self-consistency checks cannot guarantee the accuracy of photogrammetric reconstructions, a limitation that is structural rather than a tuning issue. It introduces a track‑leakage‑free hold‑out protocol that withholds a deterministic subset of images and tests each against only 3D points supported by at least two retained images, ensuring no view is evaluated against the structure it helped create. Experiments on diverse datasets show that while the protocol is well‑posed, it saturates at a confidence score of 1.00 and fails to detect coherent distortion, missing large errors that can reach over 100 m. whyItMatters":"The study highlights that hold‑out self‑validation scores, increasingly used as quality evidence for metric deliverables, may be misleading and cannot replace external survey validation."

By Behnam Asadi