arXiv Computer Vision By Daikun Liu, Teng Wang, Changyin Sun

Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation

Read the original on arXiv Computer Vision →

The paper introduces HypoDepth, an event-image monocular depth estimation framework that uses a discrete Depth Hypothesis Volume (DHV) to convert depth regression into a constrained search problem. By building a lightweight 3D cost volume between DHV features and contextual features, the method performs multi-scale correlation search for stable residual optimization, enabling efficient global-to-local refinement across resolutions. Experiments on DSEC and MVSEC show state‑of‑the‑art performance, strong zero‑shot generalization, and real‑time capability on resource‑limited devices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 19

PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

PXDepth is a monocular depth estimation model that separates global context modeling from pixel-level depth prediction. It uses a large-patch Vision Transformer to capture scene context and a pixel-space predictor with Context‑Modulated Pixel Transformer blocks to preserve high‑resolution spatial details. The approach maintains fine structures and sharp boundaries while achieving competitive global depth accuracy in zero‑shot benchmarks.

By Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
arXiv Computer Vision
Sep 3

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.

By Jung-Hee Kim, Xiaoming Liu