We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate under the assumption that pre-trained video models lack native novel view synthesis capability and enforce view alignment via camera conditioning, task-specific fine-tuning, or stepwise hard denoising guidance, often suffer from artifacts and compromised global scene consistency.
RoGe is a new end‑to‑end framework for novel view synthesis that jointly learns an implicit 3D scene representation and a video diffusion model. It eliminates the need for explicit 3D intermediates by querying the implicit scene with camera rays to produce geometric features that condition the diffusion model. Experiments on DL3DV show that RoGe surpasses reconstruction‑based, generation‑based, and hybrid baselines in image quality and temporal consistency, and ablations confirm the benefits of ray‑queried features and joint training.
By Xiaolei Lang, Ze Kang, Zehao Huang, Naiyan Wang
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views are sparse or target v...
The paper introduces 4DGS-Fixer, an iterative refinement framework that uses a video diffusion model to enhance sparse-view 4D Gaussian Splatting for dynamic scene synthesis. It first fuses multi-view depth maps into dense point clouds for better geometric initialization, then applies a pretrained video restoration model to refine rendered sequences, providing pseudo-supervision for further refinement. Experiments on a benchmark dataset show the method outperforms existing baselines, achieving nearly a 2 dB PSNR improvement.
By Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, Anyi Rao, Frank Guan
Manifold4D introduces a new denoising strategy for video re‑shooting that injects a rendered point‑cloud directly into the initial noise manifold, eliminating the need for the render to be an explicit conditioning stream during denoising. This approach allows the network to rely solely on the source video as a visual condition, improving camera‑control accuracy on the DAVIS‑Traj benchmark and Vista4D set, with significant reductions in rotation and translation errors while maintaining video fidelity. User studies confirm enhanced trajectory following and dynamic consistency, especially for large yaw amplitudes and even when the render is corrupted.
By Yongqi Mao, Zijia Dai, Zhishuo Liu, Wei Xu, Kaiwei Wang, Guotao Meng