The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.
By Jung-Hee Kim, Xiaoming Liu
arXiv:2608.20788v1 Announce Type: new
Abstract: Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or...
By Byeonggwon Lee, Sanggi Lee, Siwoo Lee, Khang Truong Giang, Soohwan Song
PXDepth is a monocular depth estimation model that separates global context modeling from pixel-level depth prediction. It uses a large-patch Vision Transformer to capture scene context and a pixel-space predictor with Context‑Modulated Pixel Transformer blocks to preserve high‑resolution spatial details. The approach maintains fine structures and sharp boundaries while achieving competitive global depth accuracy in zero‑shot benchmarks.
By Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
arXiv:2609.35734v2 Announce Type: replace
Abstract: Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, whi...
By Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang, Chunhua Shen, Mulin Yu, Bo Dai
arXiv:2609.01172v1 Announce Type: new
Abstract: Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruct...
By Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun, Yi-Hua Huang, Ziyi Yang, Peng Dai, Xiaojuan Qi
arXiv:2608.28895v1 Announce Type: new
Abstract: We introduce ReconSplat, a feed-forward model for 3D scene reconstruction that aims to address the longstanding trade-off between plausible view genera...
By Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski, Stefan Roth
GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.
By Yufei Zhang, Chenlu Zhan, Hongwei Wang
The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
arXiv:2606. 29600v1 Announce Type: cross Abstract: A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces.
By Xiaohao Xu, Feng Xue, Xiang Li, Haowei Li, Shusheng Yang, Tianyi Zhang, Matthew Johnson-Roberson, Xiaonan Huang
arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.
By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen
arXiv:2609.23404v1 Announce Type: new
Abstract: Point cloud completion aims to infer a complete 3D shape from a partial point cloud and serves as a fundamental building block for downstream tasks suc...
By Shenghui Wu, Chen Wang, Yuan Feng, Guangshun Wei, Yuanfeng Zhou, Changjian Li
PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.
By Sungjae Choi, Seunghee Koh, Junmo Kim