arXiv Computer Vision

CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving

arXiv Computer Vision
4d ago

FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

FRUC is a feedforward 3D Gaussian Splatting framework that reconstructs dynamic scenes from uncalibrated collaborative driving views. It uses a visual‑grounded geometric Transformer backbone for one‑shot, calibration‑free inference and introduces an ego‑centric causal occlusion field to model occlusion evolution across agents. The method performs cross‑agent integration as a deterministic residual denoising process, achieving state‑of‑the‑art rendering quality and efficiency on V2X‑Real and UrbanIng‑V2X datasets.

By Yihang Tao, Yu Guo, Zhengru Fang, Haonan An, Yuguang Fang
Hugging Face Trending Papers
Aug 10

Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives

Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions.

Hugging Face Trending Papers
Jun 30

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner.

arXiv Machine Learning
Jul 28

SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.

By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian
arXiv Computer Vision
Sep 25

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD introduces a multimodal representation that fuses pre‑trained 2D image and 3D LiDAR foundation models for robotic novel view synthesis, avoiding the need for a separate cross‑modal translator. By projecting LiDAR onto the image latent grid and injecting the resulting geometry‑aware packets via a lightweight residual adapter, the method enhances both RGB and depth synthesis on the GrandTour dataset compared to an image‑only baseline. Ablation studies confirm that pixel‑aligned LiDAR content drives the performance gains, and real‑world deployment on a ground robot demonstrates a tunable quality–cost trade‑off.

By Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno
arXiv Computer Vision
Oct 2

3DROID: A Renderable 3D Gaussian Dataset with Measured Per-Scene Reliability

The paper introduces 3DROID, a dataset of renderable 3D Gaussian scenes that are anchored to a robot’s metric workspace and include per-scene reliability metrics. It examines how the reliability of camera extrinsics and pose conditioning affect the fidelity of 3D Gaussian representations, proposing a calibration-aware pipeline that improves novel-view rendering when extrinsics are trustworthy. The resulting dataset, available on Hugging Face, provides robot manipulation researchers with metric-scale, pose-anchored 3D data and reliability annotations.

By Wonguen Cho, Junhoo Lee, Nojun Kwak