arXiv:2511. 15022v2 Announce Type: replace-cross Abstract: Complex-valued Gaussian primitives have recently been explored for representing holographic radiance fields in 3D novel view synthesis.
By Yicheng Zhan, Xiangjun Gao, Long Quan, Kaan Ak\c{s}it
arXiv:2606. 29600v1 Announce Type: cross Abstract: A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces.
By Xiaohao Xu, Feng Xue, Xiang Li, Haowei Li, Shusheng Yang, Tianyi Zhang, Matthew Johnson-Roberson, Xiaonan Huang
arXiv:2608. 14702v1 Announce Type: cross Abstract: Film emulation reproduces the look of an analog film stock on a new digital photograph.
By Yitong Mu
GS-PI introduces an optimization‑decoupled framework that transforms Gaussian Splatting (GS) assets into physically based rendering (PBR) compatible Gaussian assets. By treating PBR material generation as a geometry‑conditioned diffusion process on 3D point clouds, it achieves multi‑view consistency and avoids the pixel‑correspondence problems of 2D diffusion. The method employs a multi‑scale cross‑view conditioning mechanism—combining global semantic priors, photometric cues, and spatial view‑direction signals—to prevent specular highlights from baking into intrinsic colors, and then distills the predicted attributes back into a fully relightable PBR‑GS asset without requiring proxy meshes.
By Jieting Xu, Rengan Xie, Zijian Huang, Zehui Jin, Rui Wang, Yuchi Huo
arXiv:2511. 20853v4 Announce Type: replace-cross Abstract: Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent lack of large-scale, full-frame, high fidelity, real-image datasets.
By Nisarg K. Trivedi, Vinayaka A. Belludi, Li-Yun Wang
arXiv:2511. 21035v2 Announce Type: replace Abstract: Holography offers significant potential for AR/VR applications.
By Shima Rafiei, Zahra Nabizadeh Shahr-Babak, Soroush Khoubyarian, Alexandre Cooper, Shadrokh Samavi, Shahram Shirani
Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.
arXiv:2608.20788v1 Announce Type: new
Abstract: Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or...
By Byeonggwon Lee, Sanggi Lee, Siwoo Lee, Khang Truong Giang, Soohwan Song
arXiv:2609.28300v1 Announce Type: new
Abstract: Ill-posed inverse problems require priors to constrain the solution space toward plausible outcomes. In inverse rendering, learned priors modeling the...
By Andreea Ardelean, Bernhard Egger
The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.
By Jung-Hee Kim, Xiaoming Liu
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.
arXiv:2607. 12433v1 Announce Type: cross Abstract: Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE).
By Zijie Wang, Wei Zhang, Weiming Zhang, Xiao Tan, Weikai Chen, Xiaoxu Li, Guanbin Li