arXiv:2604. 03497v2 Announce Type: replace-cross Abstract: Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-crafted rewards with semantically grounded signals; however, deploying such simulation-trained policies on real vehicles remains a fundamental challenge, because they rely on simulator-native observations and simulator-coupled action semantics with no counterpart on physical hardware.
By Zilin Huang, Zhengyang Wan, Zihao Sheng, Boyue Wang, Junwei You, Sikai Chen
The paper introduces a low‑cost, open experimental platform for end‑to‑end autonomous driving on miniature Ackermann vehicles, combining a physical car, printed track, data collection tools, trajectory registration, and a Webots digital twin. It implements command‑conditioned behavior cloning, achieving a mean cross‑track error of 6.1 cm on the real vehicle and demonstrating the impact of camera field of view in simulation. Using synthetic data from the digital twin and a sim‑to‑real image translator, a higher‑capacity policy trained on both synthetic and real data completes all four track routes, outperforming the baseline trained only on real data.
By Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer
arXiv:2608. 10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes.
By Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University)
As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions dynamically update the simulator state and directly influence the next set of generated sensor observations.
arXiv:2603. 18315v2 Announce Type: replace-cross Abstract: Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the rich contextual understanding required for safe driving and make unsafe exploration unavoidable in real-world settings.
By Zilin Huang, Zihao Sheng, Zhengyang Wan, Yansong Qu, Junwei You, Sicong Jiang, Sikai Chen
The paper introduces a low‑cost, open experimental platform for end‑to‑end autonomous driving on miniature Ackermann vehicles. It combines a physical vehicle, a printed urban track, data collection tools, trajectory registration, and a Webots digital twin to link simulation and real‑world experiments. Using command‑conditioned behavior cloning, the authors demonstrate that a neural policy can follow lanes with a mean cross‑track error of 6.1 cm, and that synthetic data plus a sim‑to‑real image translator improves performance on all track routes.