The paper introduces GPERT, a framework that separates event-based 3D Gaussian Splatting into two rendering branches: event-by-event geometry rendering and snapshot-based radiance rendering. By employing ray-tracing and warped event images, GPERT balances accuracy and temporal resolution, achieving state‑of‑the‑art results on real‑world datasets and competitive performance on synthetic data. The method operates without pretrained models or COLMAP initialization, offers flexible event selection, and produces sharp reconstructions of scene edges with rapid training.
By Kai Kohyama, Yoshimitsu Aoki, Guillermo Gallego, Shintaro Shiba
arXiv:2609.39960v1 Announce Type: new
Abstract: 4D scene reconstruction aims to recover the evolving geometry, appearance, and motion of dynamic environments from visual observations. Despite substan...
By Ziren Gong, Guo Chen, Yongjia Li, Yihua Shao, Fabio Tosi, Stefano Mattoccia, Matteo Poggi, Hao Tang, Fei Ma, Shuyan Li, Ziyang Yan, Nicu Sebe, Ling Shao, Jianfei Cai, Qi Tian, Ming-Hsuan Yang
arXiv:2601.15951v2 Announce Type: replace
Abstract: Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle t...
By Sheng Miao, Sijin Li, Pan Wang, Dongfeng Bai, Bingbing Liu, Yue Wang, Andreas Geiger, Yiyi Liao
PIVOT is a new multi‑trajectory dataset and evaluation framework that captures real‑world scenes with diverse camera paths, preserving both sensor‑derived measured poses and COLMAP‑optimized poses along with calibrated and optimized intrinsics. It defines three benchmark families—seen vs. unseen trajectory generalization, measured vs. optimized pose sensitivity, and calibrated vs. optimized intrinsics sensitivity—and introduces a directed pose‑space Chamfer distance to assess pose coverage. The first version of PIVOT includes five scenes recorded with a DJI Mini 4 Pro and offers an open processing and Nerfstudio‑based evaluation toolchain, revealing a consistent quality gap between held‑out and unseen trajectories and significant sensitivity to pose source and camera intrinsics.
By Mary Raymond
arXiv:2603.14528v3 Announce Type: replace
Abstract: In recent years, 3D visual foundation models, pioneered by pointmap-based approaches such as DUSt3R, have attracted a lot of interest, achieving im...
By Shuang Guo, Filbert Febryanto, Lei Sun, Luc Van Gool, Guillermo Gallego
arXiv:2609.22500v1 Announce Type: new
Abstract: Autonomous navigation requires precise and efficient semantic segmentation, yet existing frame-based approaches remain limited by motion blur, glare, l...
By Dalia Hareb, Jean Martinet, Benoit Miramond, Elisabetta Chicca
Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and motion blur. Although neural rendering has advanced sparse-view synthesis, existing blur-aware methods typically require substantial multi-view redundancy, accurate camera poses, or costly per-scene optimization.
arXiv:2607.08098v2 Announce Type: replace
Abstract: Event cameras are increasingly adopted in embodied perception for their microsecond temporal resolution, high dynamic range, and resilience to moti...
By Linli Shi, Ruijun Zhang, Ziyun Wang
GS‑Net is a lightweight plug‑and‑play module that expands sparse Structure‑from‑Motion point clouds into dense Gaussian primitives, enabling cross‑sensor view synthesis for autonomous driving. It learns a generalizable initialization for 3D Gaussian Splatting, improving rendering quality for both interpolated and extrapolated camera viewpoints. The authors introduce CARLA‑NVS, a benchmark with 12 uniformly spaced cameras, and show that GS‑Net outperforms standard 3DGS by 2.08 dB PSNR on interpolated views and 1.86 dB on extrapolated views while being 50× faster to initialize.
By Yichen Zhang, Zihan Wang, Jiali Han, Peilin Li, Jiaxun Zhang, Jianqiang Wang, Lei He, Keqiang Li
Novel view rendering of large and complex reconstructed scenes is becoming increasingly photorealistic. However, most reconstructions remain static and lack the ambient motion that makes environments immersive.
Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic range, and low power consumption. While many studies and datasets in neuromorphic vision have focused on automotive and drone applications, human-centric daily-life scenarios remain largely underrepresented, despite their importance for developing and benchmarking event-based perception systems.
Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.