arXiv Computer Vision

FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

FRUC is a feedforward 3D Gaussian Splatting framework that reconstructs dynamic scenes from uncalibrated collaborative driving views. It uses a visual‑grounded geometric Transformer backbone for one‑shot, calibration‑free inference and introduces an ego‑centric causal occlusion field to model occlusion evolution across agents. The method performs cross‑agent integration as a deterministic residual denoising process, achieving state‑of‑the‑art rendering quality and efficiency on V2X‑Real and UrbanIng‑V2X datasets.

arXiv Computer Vision
Aug 24

GS-Net: Heterogeneous Vehicle Data Reuse via Generalizable Plug-and-Play 3DGS Module

GS‑Net is a lightweight plug‑and‑play module that expands sparse Structure‑from‑Motion point clouds into dense Gaussian primitives, enabling cross‑sensor view synthesis for autonomous driving. It learns a generalizable initialization for 3D Gaussian Splatting, improving rendering quality for both interpolated and extrapolated camera viewpoints. The authors introduce CARLA‑NVS, a benchmark with 12 uniformly spaced cameras, and show that GS‑Net outperforms standard 3DGS by 2.08 dB PSNR on interpolated views and 1.86 dB on extrapolated views while being 50× faster to initialize.

By Yichen Zhang, Zihan Wang, Jiali Han, Peilin Li, Jiaxun Zhang, Jianqiang Wang, Lei He, Keqiang Li
arXiv Computer Vision
Oct 2

FedCKA: Representation-Guided Layer Personalization for Federated 3D Perception Across Driving Domains

FedCKA introduces a Centered Kernel Alignment (CKA)-based method for federated 3D perception that dynamically balances personalization and globalization. By computing layer-wise feature similarities between local client models and a global consensus model, FedCKA generates client‑specific aggregation masks to selectively share representation‑consistent layers. Experiments on a unified multi‑domain nuScenes benchmark demonstrate that FedCKA surpasses established federated baselines, improving average NDS by 7 percentage points.

By Jolle Verhoog, Ali Burak \"Unal, Holger Caesar
arXiv Computer Vision
2d ago

Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update

arXiv:2610.07031v1 Announce Type: new Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...

By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv Computer Vision
Aug 27

PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction

PIVOT is a new multi‑trajectory dataset and evaluation framework that captures real‑world scenes with diverse camera paths, preserving both sensor‑derived measured poses and COLMAP‑optimized poses along with calibrated and optimized intrinsics. It defines three benchmark families—seen vs. unseen trajectory generalization, measured vs. optimized pose sensitivity, and calibrated vs. optimized intrinsics sensitivity—and introduces a directed pose‑space Chamfer distance to assess pose coverage. The first version of PIVOT includes five scenes recorded with a DJI Mini 4 Pro and offers an open processing and Nerfstudio‑based evaluation toolchain, revealing a consistent quality gap between held‑out and unseen trajectories and significant sensitivity to pose source and camera intrinsics.

By Mary Raymond