Sparse2comm: Towards Robust Cooperative 3D Object Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces HMS‑SCP, a hierarchical multi‑scale semantic‑aware cooperative perception framework for V2X communication. It uses a spatial importance predictor to select task‑relevant grid elements at multiple scales and maps them directly into complex‑valued symbols for joint source‑channel coding, achieving ultra‑low symbol rates and noise resilience. Experiments on OPV2V and DAIR‑V2X show that HMS‑SCP maintains high‑confidence far‑field detection with sub‑16 ms latency even under severe Rayleigh fading and extreme compression.
FedCKA introduces a Centered Kernel Alignment (CKA)-based method for federated 3D perception that dynamically balances personalization and globalization. By computing layer-wise feature similarities between local client models and a global consensus model, FedCKA generates client‑specific aggregation masks to selectively share representation‑consistent layers. Experiments on a unified multi‑domain nuScenes benchmark demonstrate that FedCKA surpasses established federated baselines, improving average NDS by 7 percentage points.
arXiv:2608. 14603v1 Announce Type: cross Abstract: Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending sensing coverage beyond occlusions and mitigating blind spots.
arXiv:2606. 09634v1 Announce Type: cross Abstract: 3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications.
VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.
3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications. Long-range detection is challenging because sensing evidence is sparse; yet this ``long-range'' scenario is routine in traffic.