arXiv:2606. 27766v1 Announce Type: cross Abstract: Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe.
By Shiqiang Gong
arXiv:2606. 06423v1 Announce Type: cross Abstract: Safety-critical traffic scenario generation is essential for evaluating autonomous driving systems under rare but high-risk interactions.
By Qi Lan, Yining Tang, Yu Shen, Yi Zhou, Yuhao Wei, Jie Li, Guofa Li
The paper introduces WM‑RMoE, a World Model‑based Risk‑aware Mixture‑of‑Experts framework for autonomous overtaking. It uses a learned latent dynamics model to perform multi‑step rollouts, evaluating cumulative risk at the trajectory level, and employs a hierarchical gating mechanism to coordinate long‑, short‑horizon, and rule‑based safety experts. A Gaussian Mixture Model preserves multimodal maneuver branches, improving robustness and preventing behavioral averaging, leading to better safety compliance, decision stability, and generalization in experiments.
By Yongzhi Liu, Sunan Zhang, Jinchang Xu, Jiawei Wang, Yushu Qiu, Chen Lv, Weichao Zhuang
arXiv:2608. 10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes.
By Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University)
arXiv:2607. 18637v1 Announce Type: cross Abstract: Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions.
By Jingzheng Li, Yufei Ge, Zhijun Chen, Qianren Mao, Zizhe Wang, Binhang Qi, Bing Li, Keyu Chen, Baochang Zhang, Xianglong Liu, Philip S Yu
arXiv:2601. 00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference.
By Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan
arXiv:2510. 09041v3 Announce Type: replace-cross Abstract: Deep reinforcement learning (DRL) has demonstrated remarkable success in developing autonomous driving policies.
By Junchao Fan, Qi Wei, Ruichen Zhang, Yang Lu, Jianhua Wang, Xiaolin Chang, Bo Ai
arXiv:2603. 18315v2 Announce Type: replace-cross Abstract: Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the rich contextual understanding required for safe driving and make unsafe exploration unavoidable in real-world settings.
By Zilin Huang, Zihao Sheng, Zhengyang Wan, Yansong Qu, Junwei You, Sicong Jiang, Sikai Chen
arXiv:2606. 16480v1 Announce Type: cross Abstract: Robots deployed in the real world must plan motions across diverse scenarios without per-scenario retuning.
By Youngjae Min, Jovin D'sa, Faizan M. Tariq, David Isele, Navid Azizan, Sangjae Bae
arXiv:2510. 02695v3 Announce Type: replace-cross Abstract: In safety-critical domains where online data collection is infeasible, offline reinforcement learning (RL) is attractive only if policies achieve high returns without catastrophic lower-tail risk.
By Kai Fukazawa, Kunal Mundada, Iman Soltani
arXiv:2607. 03903v1 Announce Type: new Abstract: Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks.
By Jiayi Guan, Tianle Zhang, Li Shen, Ruiqi Zhang, Ao Zhou, Lusong Li, Guai Chen, Mengjie Li, Alois Knoll, Xiaodong He, Changjun Jiang
arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.
By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy