arXiv:2607. 00191v1 Announce Type: cross Abstract: Collaborative-perception enables multi-robot systems to enhance situational awareness by sharing perceptual information.
By Luke Chen, Cheng-Ju Wu, David R. Martin, Qilin Ye, Pramod Khargonekar, Mohammad Abdullah Al Faruque
arXiv:2606. 01015v1 Announce Type: cross Abstract: The convergence of Artificial Intelligence, the Internet of Things, and Robotics is no longer a futuristic vision; it is rapidly becoming the foundation of real-time, intelligent, and context-aware systems.
By Ranulfo Bezerra, Satoshi Tadokoro, Kazunori Ohno
arXiv:2606. 09919v1 Announce Type: cross Abstract: Perceptual uncertainty is a central challenge for heterogeneous robot teams operating in unstructured outdoor environments, where no single viewpoint affords reliable scene understanding.
By Michal P. Podolinsky, Neel P. Bhatt, Pranay Samineni, Rohan Siva, Christian Ellis, Ufuk Topcu
Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions.
arXiv:2606. 12352v1 Announce Type: cross Abstract: Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch through a doorway to assembling structures on a construction site.
By Ria Doshi, Tian Gao, Annie Chen, Chelsea Finn, Jeannette Bohg
arXiv:2603.19308v2 Announce Type: replace-cross
Abstract: In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key...
By Wentao Wang, Haoran Xu, Guang Tan
arXiv:2503. 22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition.
By Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding
D3ARC is an asynchronous distributed hierarchical framework designed for time‑critical wildfire detection using multiple robotic agents. It enables cooperative perception, shared situational awareness, and coordinated actions while a remote controller asynchronously directs each robot’s motion. The system incorporates safe navigation, coverage efficiency, and a forward‑looking capability to evaluate candidate strategies before execution, achieving up to 94% mission success and 89.4% detection confidence in realistic simulations.
By Nikolaos Koursioumpas, Lina Magoula, Nancy Alonistioti, Ramin Khalili
AeroWeaver is a new embodied‑agent harness that integrates large language model (LLM) decision making with the executable skills of individual UAVs, enabling distributed, adaptive swarm execution. It connects semantic mission decisions to governed skills, organizes role‑conditioned local agents for coordination, and refines skill selection online using role‑indexed state‑action‑reward experience. Experiments demonstrate that AeroWeaver maintains valid skill execution without a central joint‑action generator and supports reward‑guided, training‑free adaptive learning from accumulated execution experience.
By Jiabin Lou, Yirong Yang, Haopeng Wang, Xuxin Lv, Xinyu Liu, Diyuan Hou, Xuehong Liu, Rongye Shi, Wenjun Wu
arXiv:2606. 04072v1 Announce Type: cross Abstract: Deep learning models are increasingly central to autonomous vehicle (AV) pipelines, yet their integration has traditionally followed a monolithic design where perception, planning, and control execute on a single onboard computer.
By Pragya Sharma, Brian Wang, Mani Srivastava
LiteViLNet is a lightweight RGB‑geometry fusion network for road segmentation that uses a MobileNetV3 RGB encoder and a tiny depth‑wise‑separable geometry encoder. Its multi‑scale fusion module enhances modality‑specific features, performs cross‑modal interaction, and applies adaptive gating, while a depth‑wise large‑kernel bridge expands contextual support with minimal overhead. The U‑Net‑style decoder is trained with deep supervision, achieving state‑of‑the‑art performance on KITTI and ORFD benchmarks and running at up to 68.73 FPS on a Jetson Orin NX with TensorRT FP16.
By Daojie Peng, Bingtao Wang, Fulong Ma, Liang Zhang, Jun Ma
arXiv:2609.00814v1 Announce Type: new
Abstract: Remote sensing visual models have continuously advanced various interpretation tasks. However, the research process behind model improvement still heav...
By Kaiyue Kang, Qixuan He, Peijin Wang, Yingchao Feng, Chao Ren, Kangxin Wang, Wenhui Diao, Yixiao Wang, Liangjin Zhao, Kaiwen Wei, Nayu Liu, Xian Sun