arXiv AI By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

Read the original on arXiv AI →

arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models

DepthEvidence is a 4B multimodal language model that integrates dense metric depth predictions into language generation. It employs a camera‑conditioned decoder to produce full‑resolution depth maps and a dense‑to‑language interface that converts these predictions into object‑aligned geometry tokens. The model is trained with geometric supervision and instruction tuning, and it sets new state‑of‑the‑art results on a Depth‑VQA benchmark and on instance‑level metric depth estimation across nine datasets.

By Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang
arXiv Computer Vision
Sep 3

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.

By Jung-Hee Kim, Xiaoming Liu