arXiv Computer Vision

Pumpire: Unified Benchmark for Metric Distance Estimation

Hugging Face Trending Papers
Jun 1

Honey, I Shrunk the Arc de Triomphe!

Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from a persistent ''scale-collapse'' phenomenon: distant landmarks and vast landscapes are metrically underestimated. We hypothesize that this performance gap stems from a training data bottleneck, where existing metric-scale datasets are hardware-constrained to homogenous vehicle-captured LiDAR or short-range indoor scans, or consist of synthetic data that lacks the semantic complexity of the physical world.

arXiv AI
Jul 7

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.

By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen
arXiv Computer Vision
Sep 4

VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

VI3 is a model‑agnostic framework that grounds pretrained 3D foundation models (3DFMs) by using inertial measurement unit (IMU) data to provide metric scale. It initializes and preintegrates IMU readings to create a metric motion reference, which is then used to recover the scale of 3DFM outputs. The approach includes adaptable anchoring strategies for different 3DFM architectures and demonstrates scale recovery on synthetic and real aerial datasets without ground‑truth supervision.

By Ernesto Lozano, Alberto Jaenal, Javier Civera
arXiv Computer Vision
Sep 3

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.

By Jung-Hee Kim, Xiaoming Liu
arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv Computer Vision
Sep 25

FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors

FounRef is a training‑free method that refines frozen monocular foundation priors into dense metric depth by aligning them with sparse metric anchors. It validates anchors against the prior’s predictions, rejects misaligned ones, and applies a structure‑preserving solver to correct depth globally and locally while preserving fine geometry. The approach works out of the box on unseen cameras and scenes, achieving up to 24% lower depth error, 92% lower surface‑normal noise, and nearly 15× faster inference than a leading depth‑completion network.

By Dan Halperin, Mirko M\"ahlisch