PXDepth is a monocular depth estimation model that separates global context modeling from pixel-level depth prediction. It uses a large-patch Vision Transformer to capture scene context and a pixel-space predictor with Context‑Modulated Pixel Transformer blocks to preserve high‑resolution spatial details. The approach maintains fine structures and sharp boundaries while achieving competitive global depth accuracy in zero‑shot benchmarks.
By Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.
arXiv:2609.23796v1 Announce Type: new
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challe...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
arXiv:2609.23796v2 Announce Type: replace
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open ch...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
arXiv:2609.35734v2 Announce Type: replace
Abstract: Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, whi...
By Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang, Chunhua Shen, Mulin Yu, Bo Dai
arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.
By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen
arXiv:2608.20788v1 Announce Type: new
Abstract: Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or...
By Byeonggwon Lee, Sanggi Lee, Siwoo Lee, Khang Truong Giang, Soohwan Song
The paper introduces the Geometry‑Native Autoencoder (GAE), a compact latent space that can be decoded into appearance, depth, camera parameters, and point maps, enabling 3D‑consistent world generation. By reparameterizing a geometry foundation model’s features, GAE replaces traditional appearance‑centric latents and improves visual quality and 3D coherence, achieving significant reductions in FVD and camera‑trajectory error on benchmark datasets. The work demonstrates that a geometry‑native latent space can serve as a shared interface between perception and generation models.
By Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
The paper introduces NAIMA, a guided depth super‑resolution framework that leverages global contextual semantic priors from pretrained vision transformer token embeddings. Its Guided Token Attention (GTA) module uses depth encodings as queries to attend over semantic tokens, with a zero‑initialized gate controlling the influence of semantic evidence. NAIMA achieves competitive in‑distribution performance while delivering superior cross‑dataset generalization without relying on decoded priors or auxiliary objectives.
By Tayyab Nasir, Daochang Liu, Ajmal Mian
arXiv:2608.28895v1 Announce Type: new
Abstract: We introduce ReconSplat, a feed-forward model for 3D scene reconstruction that aims to address the longstanding trade-off between plausible view genera...
By Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski, Stefan Roth
DepthEvidence is a 4B multimodal language model that integrates dense metric depth predictions into language generation. It employs a camera‑conditioned decoder to produce full‑resolution depth maps and a dense‑to‑language interface that converts these predictions into object‑aligned geometry tokens. The model is trained with geometric supervision and instruction tuning, and it sets new state‑of‑the‑art results on a Depth‑VQA benchmark and on instance‑level metric depth estimation across nine datasets.
By Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang