Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions.
arXiv:2609.16852v1 Announce Type: new
Abstract: Industrial IoT environments increasingly deploy autonomous mobile robots for tasks such as material handling, product assembly, or infrastructure inspe...
By Houssam Hajj Hassan (L2S), Antonia Maria Masucci (L2S), Lynda Zitoune (L2S), Salah-Eddine Elayoubi (L2S)
arXiv:2309.10164v3 Announce Type: replace-cross
Abstract: We develop a decentralized Perception-Action-Communication (PAC) system for multi-robot teams that enables them to collaborate in large scale...
By Saurav Agarwal, Frederic Vatnsdal, Romina Garcia Camargo, Carlos Nieto-Granda, Vijay Kumar, Alejandro Ribeiro
Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner.
arXiv:2609.00951v1 Announce Type: new
Abstract: Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability o...
By Jiuwu Hao, Ziyi Ni, Liguo Sun, Yuting Wan, Yueyang Wu, Ti Xiang, Haolin Song, Pin Lv
arXiv:2606. 12352v1 Announce Type: cross Abstract: Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch through a doorway to assembling structures on a construction site.
By Ria Doshi, Tian Gao, Annie Chen, Chelsea Finn, Jeannette Bohg
arXiv:2508. 00917v2 Announce Type: replace-cross Abstract: Connected autonomous vehicles (CAVs) must simultaneously perform multiple tasks, such as perception, prediction, planning, and control, to ensure safe and reliable navigation in complex environments.
By Jiayuan Wang, Farhad Pourpanah, Q. M. Jonathan Wu, Ning Zhang
Air-Ground Collaborative Vision-and-Language Navigation (AGC-VLN) pairs an unmanned aerial vehicle (UAV) with a global bird’s‑eye view and an unmanned ground vehicle (UGV) with a local first‑person view, creating a shared bird’s‑eye map that enables collaboration. The training‑free baseline decomposes navigation into VLM‑based semantic reasoning and deterministic geometric execution, allowing the UAV to render the UGV’s pose and target markers while the UGV plans a road‑following path using the shared map. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and surpasses the strongest single‑agent baseline by 24.0 points.
Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability of individual agents. However, most existing metho...
HeteroPROMPT is a real‑time, privacy‑preserving framework for heterogeneous collaborative perception in autonomous systems. It aligns features from diverse sensors and models into a unified ego‑centric space using modular prompts and lightweight tuning, while keeping encoders and fusion stacks frozen. The system employs a metadata‑free autoencoder for modality classification and routing, achieving higher average precision on OPV2V‑H and V2XSet datasets with far fewer trainable parameters.
By Armin Maleki, Hayder Radha
PEARL is a lightweight, prompt‑embedding framework designed for real‑time, anonymous, and heterogeneous collaborative perception. It uses two parallel, low‑rank visual prompt interpreters—sparse‑detection (LWSD) and dense, domain‑invariant (LWDDI)—to align features and select the appropriate interpreter for newly joining agents without needing their configurations. Experiments on simulated and real datasets show that PEARL improves average precision by 8.2% over random selection, runs in 1.67 ms, and reduces communication cost by up to 34.7× while outperforming state‑of‑the‑art offline methods by 5.6% AP.
By Armin Maleki, Hayder Radha
Air-Ground Collaborative Vision-and-Language Navigation (AGC‑VLN) pairs a UAV with a global bird’s‑eye view and a UGV with a local first‑person view, creating a shared bird’s‑eye map that displays the UGV’s pose and the target location. The UAV localizes the target in its downward view and flies toward it, while the UGV uses the shared map to plan a road‑following path with a frozen VLM and execute it under closed‑loop control. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and outperforms the strongest single‑agent baseline by 24.0 points.
By Shuning Zhang, Liang Li, Yunheng Wang, Tao Wang, Yihang Kang, Renjing Xu