Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not directly yield reliable 3D assets: camera motion is difficult to control, view coverage is partial, and frames often contain inconsistencies across time.
arXiv:2608.22465v1 Announce Type: new
Abstract: High-fidelity free-viewpoint video (FVV) and interactive rendering increasingly rely on explicit Gaussian representations, yet practical deployment rem...
By Xinhui Liu, Lei Liu, Zhenghao Chen, Lebin Zhou, Wei Wang, Wei Jiang
GS‑VLA introduces a lightweight, plug‑and‑play framework that uses a 4 M‑parameter 3D‑Gaussian canonicalizer to adapt frozen Vision‑Language‑Action (VLA) policies to viewpoint shifts without retraining the policy. By treating viewpoint changes as a localized novel‑view synthesis problem under a locality assumption, the method normalizes observations through a scene‑ and policy‑independent disocclusion task. Experiments on the LIBERO benchmark demonstrate that GS‑VLA recovers a large portion of performance lost due to camera displacement, improving results across different policy architectures, unseen task suites, and perturbation scales.
whyItMatters":"The approach offers a computationally efficient alternative to costly fine‑tuning or generative augmentation, enabling robust VLA deployment in real‑world settings where camera configurations may vary."
By Yechan Park, HyunJin Kim
arXiv:2511.16030v3 Announce Type: replace
Abstract: 3D Gaussian Splatting (3DGS) enables efficient, high-fidelity novel view synthesis, yet its performance degrades severely under sparse-view supervi...
By Zijian Wu, Mingfeng Jiang, Zidian Lin, Ying Song, Ziqian Lu, Qun Wu, Hanjie Ma
arXiv:2609.13504v1 Announce Type: new
Abstract: Recent developments in feed-forward 3D reconstruction resulted in models which can recover dense scene representations and camera motion solely from an...
By Tingjun Huang, Dmitry Rudshin, Mathieu Meyer, Pietro Bonazzi, Marc Pollefeys, Emilia Szyma\'nska
RoGe is a new end‑to‑end framework for novel view synthesis that jointly learns an implicit 3D scene representation and a video diffusion model. It eliminates the need for explicit 3D intermediates by querying the implicit scene with camera rays to produce geometric features that condition the diffusion model. Experiments on DL3DV show that RoGe surpasses reconstruction‑based, generation‑based, and hybrid baselines in image quality and temporal consistency, and ablations confirm the benefits of ray‑queried features and joint training.
By Xiaolei Lang, Ze Kang, Zehao Huang, Naiyan Wang