ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
arXiv:2607. 12433v1 Announce Type: cross Abstract: Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE).
arXiv:2607. 12433v1 Announce Type: cross Abstract: Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE).
arXiv:2606. 29600v1 Announce Type: cross Abstract: A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces.
PXDepth is a monocular depth estimation model that separates global context modeling from pixel-level depth prediction. It uses a large-patch Vision Transformer to capture scene context and a pixel-space predictor with Context‑Modulated Pixel Transformer blocks to preserve high‑resolution spatial details. The approach maintains fine structures and sharp boundaries while achieving competitive global depth accuracy in zero‑shot benchmarks.
Dual-pixel (DP) imaging enables metric depth estimation from a single camera using sub-aperture disparity. However, the extremely small effective baseline limits disparity observability, leading to structural degradation and depth failure in textureless, low-contrast, or downsampled regions.
arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.
3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view settings due to insufficient geometric constraints. While recent methods introduce monocular depth priors to mitigate this, they inherently struggle with scale ambiguity and cross-view inconsistency, leading to defective geometry.
arXiv:2608.21136v1 Announce Type: new Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...
arXiv:2608.22821v1 Announce Type: new Abstract: We present SiZeUp, a fast and scalable approach for constructing large-scale 3D urban proxy models directly from calibrated oblique aerial imagery. Our...
arXiv:2604. 05182v2 Announce Type: replace-cross Abstract: We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.
arXiv:2512. 16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving.
arXiv:2607. 16286v1 Announce Type: cross Abstract: The 3D geometry of real-world scene data is often incomplete.
arXiv:2607. 16012v1 Announce Type: cross Abstract: Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation.