arXiv Computer Vision

Sparse2comm: Towards Robust Cooperative 3D Object Detection

arXiv AI
Aug 19

HMS-SCP: Task-Oriented Multi-Scale Semantic Communication for V2X Cooperative Perception

The paper introduces HMS‑SCP, a hierarchical multi‑scale semantic‑aware cooperative perception framework for V2X communication. It uses a spatial importance predictor to select task‑relevant grid elements at multiple scales and maps them directly into complex‑valued symbols for joint source‑channel coding, achieving ultra‑low symbol rates and noise resilience. Experiments on OPV2V and DAIR‑V2X show that HMS‑SCP maintains high‑confidence far‑field detection with sub‑16 ms latency even under severe Rayleigh fading and extreme compression.

By Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim
arXiv Computer Vision
Oct 2

FedCKA: Representation-Guided Layer Personalization for Federated 3D Perception Across Driving Domains

FedCKA introduces a Centered Kernel Alignment (CKA)-based method for federated 3D perception that dynamically balances personalization and globalization. By computing layer-wise feature similarities between local client models and a global consensus model, FedCKA generates client‑specific aggregation masks to selectively share representation‑consistent layers. Experiments on a unified multi‑domain nuScenes benchmark demonstrate that FedCKA surpasses established federated baselines, improving average NDS by 7 percentage points.

By Jolle Verhoog, Ali Burak \"Unal, Holger Caesar
arXiv Computer Vision
Sep 21

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.

By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao
Hugging Face Trending Papers
Jun 30

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner.

arXiv Computer Vision
1d ago

LR-V2X: Loss-Resilient Collaborative Perception under Low-Bandwidth Communication

LR‑V2X is a loss‑resilient collaborative perception framework for vehicular networks that reconstructs missing bird‑view (BEV) features from corrupted latent representations, even under severe packet loss. It converts corrupted latents into a spatial prior and uses ego‑vehicle context to recover BEV information, requiring only training under full‑communication conditions. Experiments on DAIR‑V2X and V2XREAL demonstrate that LR‑V2X maintains robust collaboration while reducing communication overhead by 64× compared to dense BEV fusion methods.

By Kang Yang, Tianci Bu, Peng Wang, Deying Li, Yongcai Wang