arXiv Computer Vision

Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception

The paper introduces a byte‑constrained cooperative perception framework that balances dense coverage with sparse refinement. Each vehicle sends a highly compressed coarse Bird’s‑Eye‑View (BEV) layer covering the entire map and uses the remaining bandwidth to transmit high‑resolution patches selected by a Task‑Aware Benefit Selector. Experiments on DAIR‑V2X and OPV2V demonstrate that this coverage‑refinement strategy achieves superior accuracy‑payload trade‑offs, reaching 0.60 AP@0.7 with only 1.87 KB per non‑ego agent.

arXiv Computer Vision
Sep 11

HeteroPROMPT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework

HeteroPROMPT is a real‑time, privacy‑preserving framework for heterogeneous collaborative perception in autonomous systems. It aligns features from diverse sensors and models into a unified ego‑centric space using modular prompts and lightweight tuning, while keeping encoders and fusion stacks frozen. The system employs a metadata‑free autoencoder for modality classification and routing, achieving higher average precision on OPV2V‑H and V2XSet datasets with far fewer trainable parameters.

By Armin Maleki, Hayder Radha
arXiv Machine Learning
Sep 24

PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception

PEARL is a lightweight, prompt‑embedding framework designed for real‑time, anonymous, and heterogeneous collaborative perception. It uses two parallel, low‑rank visual prompt interpreters—sparse‑detection (LWSD) and dense, domain‑invariant (LWDDI)—to align features and select the appropriate interpreter for newly joining agents without needing their configurations. Experiments on simulated and real datasets show that PEARL improves average precision by 8.2% over random selection, runs in 1.67 ms, and reduces communication cost by up to 34.7× while outperforming state‑of‑the‑art offline methods by 5.6% AP.

By Armin Maleki, Hayder Radha
arXiv Machine Learning
Jul 22

Recti-Q: Feature-Space Rectification for Out-of-Distribution-Robust Quantized Perception in Edge Robotics

arXiv:2607. 18540v1 Announce Type: cross Abstract: Robotic perception pipelines increasingly rely on large vision backbones deployed on SWaP-constrained edge platforms, making post-training quantization (PTQ) attractive for real-time inference.

By Hamidreza Yaghoubi Araghi, Parastoo Pilevar, Ming C. Lin
arXiv AI
Aug 19

HMS-SCP: Task-Oriented Multi-Scale Semantic Communication for V2X Cooperative Perception

The paper introduces HMS‑SCP, a hierarchical multi‑scale semantic‑aware cooperative perception framework for V2X communication. It uses a spatial importance predictor to select task‑relevant grid elements at multiple scales and maps them directly into complex‑valued symbols for joint source‑channel coding, achieving ultra‑low symbol rates and noise resilience. Experiments on OPV2V and DAIR‑V2X show that HMS‑SCP maintains high‑confidence far‑field detection with sub‑16 ms latency even under severe Rayleigh fading and extreme compression.

By Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim
arXiv Machine Learning
Jul 28

SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.

By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian