Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from a persistent ''scale-collapse'' phenomenon: distant landmarks and vast landscapes are metrically underestimated. We hypothesize that this performance gap stems from a training data bottleneck, where existing metric-scale datasets are hardware-constrained to homogenous vehicle-captured LiDAR or short-range indoor scans, or consist of synthetic data that lacks the semantic complexity of the physical world.
arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.
By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen
arXiv:2609.01172v1 Announce Type: new
Abstract: Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruct...
By Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun, Yi-Hua Huang, Ziyi Yang, Peng Dai, Xiaojuan Qi
VI3 is a model‑agnostic framework that grounds pretrained 3D foundation models (3DFMs) by using inertial measurement unit (IMU) data to provide metric scale. It initializes and preintegrates IMU readings to create a metric motion reference, which is then used to recover the scale of 3DFM outputs. The approach includes adaptable anchoring strategies for different 3DFM architectures and demonstrates scale recovery on synthetic and real aerial datasets without ground‑truth supervision.
By Ernesto Lozano, Alberto Jaenal, Javier Civera
arXiv:2608.22821v1 Announce Type: new
Abstract: We present SiZeUp, a fast and scalable approach for constructing large-scale 3D urban proxy models directly from calibrated oblique aerial imagery. Our...
By Wenjun Zhou, Yunshan Li, Qiaoyu Zhu, Weidan Xiong, Hao Zhang, Daniel Cohen-Or, Hui Huang
arXiv:2609.09394v1 Announce Type: new
Abstract: Recovering metric 3D geometry from monocular images is a fundamental computer vision task, yet current methods remain heavily fragmented by fixed camer...
By Botao Ye, Marc Pollefeys, Ming-Hsuan Yang, Abhijit Kundu
The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.
By Jung-Hee Kim, Xiaoming Liu
TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.
By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv:2603.14528v3 Announce Type: replace
Abstract: In recent years, 3D visual foundation models, pioneered by pointmap-based approaches such as DUSt3R, have attracted a lot of interest, achieving im...
By Shuang Guo, Filbert Febryanto, Lei Sun, Luc Van Gool, Guillermo Gallego
FounRef is a training‑free method that refines frozen monocular foundation priors into dense metric depth by aligning them with sparse metric anchors. It validates anchors against the prior’s predictions, rejects misaligned ones, and applies a structure‑preserving solver to correct depth globally and locally while preserving fine geometry. The approach works out of the box on unseen cameras and scenes, achieving up to 24% lower depth error, 92% lower surface‑normal noise, and nearly 15× faster inference than a leading depth‑completion network.
By Dan Halperin, Mirko M\"ahlisch
We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data design. Stage 1 learns a high-fidelity, affine-invariant geometry model by initializing with DINOv3 and training on a curated 10.
arXiv:2608.20788v1 Announce Type: new
Abstract: Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or...
By Byeonggwon Lee, Sanggi Lee, Siwoo Lee, Khang Truong Giang, Soohwan Song