VIPS is a benchmark for vehicle‑to‑infrastructure cooperative autonomous driving that uses pseudo‑simulation to combine vehicle and infrastructure observations, enabling scalable yet realistic evaluation of robustness and error propagation without full simulation. The paper also introduces CoS‑V2X, a cooperative planning framework that employs sparse representations to model vehicle‑infrastructure interactions efficiently and robustly under heterogeneous observations.
By Hoonhee Cho, Jae-Young Kang, Giwon Lee, Hyemin Yang, Heejun Park, Kuk-Jin Yoon
arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.
By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian
MILER is an end‑to‑end reinforcement learning framework that achieves zero‑shot sim‑to‑real transfer for autonomous driving in unstructured environments. It uses a custom semantic mid‑level representation (MLR) simulator for offline training, and during deployment it processes real camera and LiDAR data with BEVFusion to produce a compatible bird’s‑eye‑view representation. The policy’s actions are applied via a trajectory‑alignment strategy, allowing the system to drive 17.3 km on a 3.0 km test track without human intervention, all running on a Jetson AGX Orin.
By Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch
arXiv:2508. 00917v2 Announce Type: replace-cross Abstract: Connected autonomous vehicles (CAVs) must simultaneously perform multiple tasks, such as perception, prediction, planning, and control, to ensure safe and reliable navigation in complex environments.
By Jiayuan Wang, Farhad Pourpanah, Q. M. Jonathan Wu, Ning Zhang
FluxVLA Engine is an open, configuration‑driven platform that unifies the fragmented components of embodied policy development—datasets, visual‑language and world models, action heads, learning methods, distributed training, simulation evaluation, inference, and robot interfaces—into a reproducible data‑to‑deployment workflow. It adds features such as compositional dual‑arm simulation, scalable automatic data generation, human‑in‑the‑loop rollout and correction, Real‑Time Chunking for fast inference, and lightweight remote GPU serving, thereby linking offline learning, simulation validation, online correction, and real‑robot execution under shared, auditable contracts. The engine aims to eliminate engineering bottlenecks that currently separate promising embodied‑learning algorithms from reliable, reproducible deployment.
By Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang, Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen
arXiv:2606. 01312v1 Announce Type: cross Abstract: The integration of Artificial Intelligence (AI) and emerging 6G networks introduces new opportunities for scalable coordination in tactical autonomous vehicle systems.
By Kiran Khurshid, Shumaila Javaid, Nasir Saeed
arXiv:2609.39964v1 Announce Type: new
Abstract: Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. D...
By Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner.
arXiv:2607. 13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains.
By Zhouchonghao Wu, Akshay Rangesh, Weixin Li, Wei-Jer Chang, Zachary Lee, Tim Wang, Wei Zhan
arXiv:2607. 16074v1 Announce Type: cross Abstract: The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives.
By Haoran Sun, Wentao Zhang, Junyang Hua, Hedan Yang, Yongjian Guo, Yifei Zhang, Xiaolong Xiang, Mingxi Luo, Jing Long, Chen Zhao, Chen Zhou, Wanting Xu, Qiming Yang, Hui Zhang, Song Wang, Xiaodong Bai, Shuai Di, Xu Chu, Xiaotie Deng, Yicheng Gong, Junwu Xiong
The paper reviews end‑to‑end autonomous driving (E2E‑AD) training, framing it as a Data‑Strategy‑Platform system. It surveys recent advances in data pipelines, learning paradigms, and training infrastructures, and discusses how these layers interact to influence model performance, robustness, and deployability. The authors highlight current limitations and propose a future vision that prioritizes data value, foundation‑driven generalization, and integrated training‑testing loops for more robust, scalable, and trustworthy autonomous driving systems.
By Chengkai Xu, Yiming Cui, Jiaqi Liu, Yicheng Guo, Cheng Qin, Geyuan Zhang, Xinwei Dong, Shiyu Fang, Peng Hang, Jian Sun
The paper introduces CloudEdgeVLA, a cloud‑edge policy for Vision‑Language‑Action models that treats temporal misalignment as a representation‑learning problem. It encodes delayed observations into slowly varying task features on the cloud while a lightweight edge head fuses the latest cloud feature with current local vision. Experiments on four LIBERO suites show that CloudEdgeVLA retains 63.8–78.0% success under a 40‑step delay window, far outperforming VLASH and single‑path baselines.
By Daojie Peng, Fulong Ma, Bingtao Wang, Sheng Wang, Jun Ma