HeteroPROMPT is a real‑time, privacy‑preserving framework for heterogeneous collaborative perception in autonomous systems. It aligns features from diverse sensors and models into a unified ego‑centric space using modular prompts and lightweight tuning, while keeping encoders and fusion stacks frozen. The system employs a metadata‑free autoencoder for modality classification and routing, achieving higher average precision on OPV2V‑H and V2XSet datasets with far fewer trainable parameters.
By Armin Maleki, Hayder Radha
arXiv:2603.19308v2 Announce Type: replace-cross
Abstract: In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key...
By Wentao Wang, Haoran Xu, Guang Tan
arXiv:2609.00951v1 Announce Type: new
Abstract: Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability o...
By Jiuwu Hao, Ziyi Ni, Liguo Sun, Yuting Wan, Yueyang Wu, Ti Xiang, Haolin Song, Pin Lv
BOLT is a lightweight plug‑and‑play module that enables preparation‑free heterogeneous cooperative perception by adapting neighboring features online through ego‑as‑teacher distillation. It requires only ego predictions, no ground‑truth labels, and uses high‑confidence ego features to align cross‑agent feature domains while allowing neighbors to contribute in low‑confidence regions. With just 0.9 M trainable parameters, BOLT boosts AP@50 by up to 32.3 points over unadapted fusion and consistently outperforms ego‑only results on DAIR‑V2X and OPV2V.
By Kang Yang, Tianci Bu, Peng Wang, Deying Li, Yongcai Wang
Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability of individual agents. However, most existing metho...
arXiv:2608. 07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments.
By Peng Xu, Chengcheng Wang, Shaohua Wan
The paper introduces LDE, a framework that uses collaborative perception (CP) to generate high‑quality pseudo‑labels for unsupervised model adaptation in autonomous driving. It tackles communication limits, field‑of‑view mismatches, and label unreliability through selective feature sharing, FoV filtering, and curriculum learning. Experiments on 3D object detection show LDE surpasses pre‑trained models and existing adaptation methods.
By Yanan Ma, Yihang Tao, Zhengru Fang, Zihan Fang, Yiqin Deng, Xianhao Chen, Yuguang Fang
arXiv:2609.17856v1 Announce Type: new
Abstract: Heterogeneous cooperative perception (CP) enables connected vehicles with diverse sensor setups to share spatial awareness via compact feature maps, wh...
By Chenyi Wang, Yutong Liu, Qingzhao Zhang, Ming F. Li
VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.
By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao
arXiv:2602. 14401v2 Announce Type: replace-cross Abstract: Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant privacy concerns.
By Qingqian Yang, Hao Wang, Sai Qian Zhang, Jian Li, Yang Hua, Miao Pan, Tao Song, Zhengwei Qi, Haibing Guan
arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.
By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian
The paper introduces a byte‑constrained cooperative perception framework that balances dense coverage with sparse refinement. Each vehicle sends a highly compressed coarse Bird’s‑Eye‑View (BEV) layer covering the entire map and uses the remaining bandwidth to transmit high‑resolution patches selected by a Task‑Aware Benefit Selector. Experiments on DAIR‑V2X and OPV2V demonstrate that this coverage‑refinement strategy achieves superior accuracy‑payload trade‑offs, reaching 0.60 AP@0.7 with only 1.87 KB per non‑ego agent.
By Melih Yazgan, Timon M\"uller, J. Marius Z\"ollner