The paper introduces Ref-GeNVS, a training‑free, reflection‑aware approach for generative novel view synthesis in mirror scenes. It treats a mirror image as two complementary views, estimates the mirror plane and reflected camera poses, and uses a two‑stage generation process with Mirror‑gated attention and Reflection injection to produce reflection‑consistent novel views. The method leverages a multi‑view diffusion backbone without finetuning, outperforming recent generative NVS methods on synthetic and real mirror scenes.
By GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh
arXiv:2608. 07463v1 Announce Type: cross Abstract: Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis.
By Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau
arXiv:2608. 07559v1 Announce Type: cross Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models.
By Yidan Shen, Yu Wen, Chen Zhang, Xin Fu, Renjie Hu
arXiv:2605.12957v2 Announce Type: replace
Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of do...
By Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen
SpatialCrafter introduces a two‑stage framework for single‑image world modeling that first generates a global 3D proxy using a Point‑anchored Sparse Structure Flow module, then refines appearance with a Generative Deferred Refiner built on a video diffusion model. The method incorporates Parallel Geometry Injection and Proxy‑Aware Corruption training to integrate the proxy without disrupting the pretrained generative manifold, and it is evaluated on a newly constructed dataset of 115K scenes. Experiments demonstrate that SpatialCrafter outperforms existing approaches, reducing long‑term drift and maintaining consistency under rapid camera motion and extreme viewpoints.
By Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan
arXiv:2608. 20212v1 Announce Type: new Abstract: High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections.
By Radim Spetlik, David Futschik, Radek Danecek, Feitong Tan, Ziqian Bai, Rohit Pandey, Yinda Zhang
arXiv:2609.10531v1 Announce Type: new
Abstract: Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the avai...
By Jerred Chen, Simon Weber, Ronald Clark
arXiv:2608. 20107v1 Announce Type: new Abstract: Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting.
By Yigit Ekin, Enes Sanli, Aykut Erdem, Erkut Erdem, Aysegul Dundar
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content creation. Recent methods achieve outstanding performance by leveraging camera-controlled video diffusion models, but rely on iterative diffusion sampling, which greatly limits their practical deployment.
The paper introduces a benchmark called Shedding Light to evaluate how well generative image models understand and reproduce lighting. The benchmark tests models by asking them to inpaint a simple object, called a light probe, into real photographs and then compares the generated probe to the ground truth to assess lighting direction, colour, and radiance. The authors provide a scalable protocol and open-source code and data for systematic assessment of photometric accuracy in future models.
By Justine Giroux, Jack Oliver Hilliard, Yannick Hold-Geoffroy, Javier Vazquez-Corral, Jean-Fran\c{c}ois Lalonde