arXiv:2608.21055v1 Announce Type: cross
Abstract: Collaborative perception extends the sensing range of a single vehicle by fusing observations from nearby agents, which improves the robustness of au...
By Chi Li, Rui Lin, Aobo Ji, Dongzhu Xu
arXiv:2610.00319v1 Announce Type: new
Abstract: Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating o...
By Lingzhao Kong, Yongsheng Zang, Yu Kang, Kailun Yang, Jie Fu, Yukun Zuo, Zhiyong Li
Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions.
arXiv:2610.08573v1 Announce Type: new
Abstract: Cooperative perception improves autonomous driving by sharing complementary observations among vehicles and roadside infrastructure for 3D object detec...
By Lei Yang, Boqi Li, Chunmian Lin, Li Wang, Ziying Song, Shaoqing Xu, Heye Huang, Haibao Yu, Chen Lv
PEARL is a lightweight, prompt‑embedding framework designed for real‑time, anonymous, and heterogeneous collaborative perception. It uses two parallel, low‑rank visual prompt interpreters—sparse‑detection (LWSD) and dense, domain‑invariant (LWDDI)—to align features and select the appropriate interpreter for newly joining agents without needing their configurations. Experiments on simulated and real datasets show that PEARL improves average precision by 8.2% over random selection, runs in 1.67 ms, and reduces communication cost by up to 34.7× while outperforming state‑of‑the‑art offline methods by 5.6% AP.
By Armin Maleki, Hayder Radha
The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.
By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
LR‑V2X is a loss‑resilient collaborative perception framework for vehicular networks that reconstructs missing bird‑view (BEV) features from corrupted latent representations, even under severe packet loss. It converts corrupted latents into a spatial prior and uses ego‑vehicle context to recover BEV information, requiring only training under full‑communication conditions. Experiments on DAIR‑V2X and V2XREAL demonstrate that LR‑V2X maintains robust collaboration while reducing communication overhead by 64× compared to dense BEV fusion methods.
By Kang Yang, Tianci Bu, Peng Wang, Deying Li, Yongcai Wang
arXiv:2609.00951v1 Announce Type: new
Abstract: Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability o...
By Jiuwu Hao, Ziyi Ni, Liguo Sun, Yuting Wan, Yueyang Wu, Ti Xiang, Haolin Song, Pin Lv
arXiv:2605. 20301v2 Announce Type: replace-cross Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making.
By Wenxuan Li, Qin Zou, Shoubing Chen, Chi Chen, Yingyi Yang, Qingxiang Meng
arXiv:2603.19308v2 Announce Type: replace-cross
Abstract: In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key...
By Wentao Wang, Haoran Xu, Guang Tan
The paper introduces OccLinker, a lightweight plugin for vision‑based occupancy networks that reduces flickering by efficiently merging historical static and motion cues with current features via a dual cross‑attention mechanism. It generates correction components to refine base network predictions and proposes a new temporal consistency metric to quantify flickering. Experiments on two benchmark datasets show that OccLinker improves performance with minimal computational overhead while effectively diminishing flickering artifacts.
By Fengcheng Yu, Haoran Xu, Canming Xia, Ziyang Zong, Guang Tan
Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability of individual agents. However, most existing metho...