We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environment, while interactive roaming requires a complete surrounding world, persistent geometry, controllable camera motion, and temporally coherent high-fidelity observations.
arXiv:2609.15032v1 Announce Type: new
Abstract: Live free-viewpoint visualization of real humans is critical for immersive communication and interactive digital experiences. Existing methods either r...
By Hanzhang Tu, Zhanfeng Liao, Wei Min, Jiajun Zhang, Yebin Liu
RoGe is a new end‑to‑end framework for novel view synthesis that jointly learns an implicit 3D scene representation and a video diffusion model. It eliminates the need for explicit 3D intermediates by querying the implicit scene with camera rays to produce geometric features that condition the diffusion model. Experiments on DL3DV show that RoGe surpasses reconstruction‑based, generation‑based, and hybrid baselines in image quality and temporal consistency, and ablations confirm the benefits of ray‑queried features and joint training.
By Xiaolei Lang, Ze Kang, Zehao Huang, Naiyan Wang
arXiv:2609.14462v1 Announce Type: new
Abstract: Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Exist...
By Jiaming Tan, Mingliang Zhai, Zhen Li, Yuwei Wu, Chuanhao Li, Kaipeng Zhang
arXiv:2609.23796v2 Announce Type: replace
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open ch...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.
By Yufei Zhang, Chenlu Zhan, Hongwei Wang
arXiv:2609.23796v1 Announce Type: new
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challe...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
MUGEN is a large‑scale real‑world dataset of over 1,300 hours of 4K panoramic videos with rich semantic and geometric annotations, designed to support interactive 360° world exploration. Wan360 is a camera‑controllable panoramic video generation model built on MUGEN, featuring ERP‑aware components (periodic longitude RoPE, ERP‑aware padding, random roll yaw) and a panoramic Plücker embedding for camera motion. Together, they address gaps in data and model support for immersive, temporally coherent 360° video generation along user‑specified camera trajectories.
By Jiaming Tan, Zhen Li, Shuwei Shi, Minggui Teng, Siqi Yang, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang
Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and motion blur. Although neural rendering has advanced sparse-view synthesis, existing blur-aware methods typically require substantial multi-view redundancy, accurate camera poses, or costly per-scene optimization.
UniQueR is a unified query‑based feedforward framework that reconstructs 3D scenes from unposed images by treating reconstruction as a sparse 3D query inference problem. It learns a compact set of 3D anchor points that serve as explicit geometric queries, allowing the network to infer scene structure—including occluded geometry—in a single forward pass. By encoding spatial and appearance priors directly in global 3D space and using a decoupled cross‑attention design, UniQueR achieves strong geometric expressiveness while reducing memory and computational cost, outperforming state‑of‑the‑art feedforward methods on Mip‑NeRF 360 and VR‑NeRF with far fewer primitives.
By Chensheng Peng, Quentin Herau, Jiezhi Yang, Yichen Xie, Yihan Hu, Wenzhao Zheng, Matthew Strong, Masayoshi Tomizuka, Wei Zhan
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.
By Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang