arXiv Machine Learning

PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception

PEARL is a lightweight, prompt‑embedding framework designed for real‑time, anonymous, and heterogeneous collaborative perception. It uses two parallel, low‑rank visual prompt interpreters—sparse‑detection (LWSD) and dense, domain‑invariant (LWDDI)—to align features and select the appropriate interpreter for newly joining agents without needing their configurations. Experiments on simulated and real datasets show that PEARL improves average precision by 8.2% over random selection, runs in 1.67 ms, and reduces communication cost by up to 34.7× while outperforming state‑of‑the‑art offline methods by 5.6% AP.

arXiv Computer Vision
Sep 11

HeteroPROMPT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework

HeteroPROMPT is a real‑time, privacy‑preserving framework for heterogeneous collaborative perception in autonomous systems. It aligns features from diverse sensors and models into a unified ego‑centric space using modular prompts and lightweight tuning, while keeping encoders and fusion stacks frozen. The system employs a metadata‑free autoencoder for modality classification and routing, achieving higher average precision on OPV2V‑H and V2XSet datasets with far fewer trainable parameters.

By Armin Maleki, Hayder Radha
arXiv Computer Vision
Sep 3

BOLT: Online Lightweight Adaptation for Preparation-Free Heterogeneous Cooperative Perception

BOLT is a lightweight plug‑and‑play module that enables preparation‑free heterogeneous cooperative perception by adapting neighboring features online through ego‑as‑teacher distillation. It requires only ego predictions, no ground‑truth labels, and uses high‑confidence ego features to align cross‑agent feature domains while allowing neighbors to contribute in low‑confidence regions. With just 0.9 M trainable parameters, BOLT boosts AP@50 by up to 32.3 points over unadapted fusion and consistently outperforms ego‑only results on DAIR‑V2X and OPV2V.

By Kang Yang, Tianci Bu, Peng Wang, Deying Li, Yongcai Wang
arXiv Machine Learning
Sep 17

Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation

The paper introduces LDE, a framework that uses collaborative perception (CP) to generate high‑quality pseudo‑labels for unsupervised model adaptation in autonomous driving. It tackles communication limits, field‑of‑view mismatches, and label unreliability through selective feature sharing, FoV filtering, and curriculum learning. Experiments on 3D object detection show LDE surpasses pre‑trained models and existing adaptation methods.

By Yanan Ma, Yihang Tao, Zhengru Fang, Zihan Fang, Yiqin Deng, Xianhao Chen, Yuguang Fang
arXiv Computer Vision
Sep 21

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.

By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao
arXiv Machine Learning
Jul 28

SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.

By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian
arXiv Computer Vision
Sep 25

Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception

The paper introduces a byte‑constrained cooperative perception framework that balances dense coverage with sparse refinement. Each vehicle sends a highly compressed coarse Bird’s‑Eye‑View (BEV) layer covering the entire map and uses the remaining bandwidth to transmit high‑resolution patches selected by a Task‑Aware Benefit Selector. Experiments on DAIR‑V2X and OPV2V demonstrate that this coverage‑refinement strategy achieves superior accuracy‑payload trade‑offs, reaching 0.60 AP@0.7 with only 1.87 KB per non‑ego agent.

By Melih Yazgan, Timon M\"uller, J. Marius Z\"ollner