M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.
arXiv:2609.01172v1 Announce Type: new Abstract: Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruct...
arXiv:2607. 12433v1 Announce Type: cross Abstract: Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE).
The paper introduces HypoDepth, an event-image monocular depth estimation framework that uses a discrete Depth Hypothesis Volume (DHV) to convert depth regression into a constrained search problem. By building a lightweight 3D cost volume between DHV features and contextual features, the method performs multi-scale correlation search for stable residual optimization, enabling efficient global-to-local refinement across resolutions. Experiments on DSEC and MVSEC show state‑of‑the‑art performance, strong zero‑shot generalization, and real‑time capability on resource‑limited devices.
The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.
FounRef is a training‑free method that refines frozen monocular foundation priors into dense metric depth by aligning them with sparse metric anchors. It validates anchors against the prior’s predictions, rejects misaligned ones, and applies a structure‑preserving solver to correct depth globally and locally while preserving fine geometry. The approach works out of the box on unseen cameras and scenes, achieving up to 24% lower depth error, 92% lower surface‑normal noise, and nearly 15× faster inference than a leading depth‑completion network.