OPTED is a method for on‑policy fine‑tuning of end‑to‑end driving models that separates reinforcement learning from the policy update. A privileged teacher trained with RL on vectorized inputs (HD‑maps and bounding boxes) supervises the pre‑trained student during closed‑loop post‑training. Applied to the camera‑based models TransFuser and VaVAM in AlpaSim, OPTED boosts driving scores by 1.6× and 9.5×, respectively, while requiring roughly three orders of magnitude fewer simulator interactions than direct RL post‑training.
By Damiano Da Col, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, Christos Sakaridis
The paper proposes a new training dataset that generates more informative positive and negative samples for trajectory scoring in autonomous driving. By perturbing logged human trajectories laterally toward the drivable boundary and longitudinally toward a leading vehicle, the dataset provides richer supervision than the planner’s default proposal pool. Using a transformer-based scorer trained on this dataset, the authors achieve improved EPDMS scores on two frozen planners, DiffusionDrive and MeanFuser, when evaluated on the NAVSIM navtrain dataset.
By Yaguang Li, Jiaru Zhang, Chuheng Wei, Can Cui, Ziran Wang
arXiv:2606. 31106v1 Announce Type: cross Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics.
By Hyeonchang Jeon, Kyungbeom Kim, Eugene Vinitsky, Kyung-Joong Kim
arXiv:2607. 13319v1 Announce Type: cross Abstract: High-speed off-road autonomy requires precise closed-loop control for a target vehicle while remaining robust across changing terrains.
By Rwik Rana, Jesse Quattrociocchi, Christian Ellis, Nathan Tsoi, Garrett Warnell, Joydeep Biswas
The paper presents a single pretrained diffusion traffic model that serves both as an ego motion planner and as a controllable generator of safety‑critical scenarios for autonomous driving. It introduces a Single‑Stream Dual‑Stream diffusion‑transformer decoder (SSDS) that fuses scene context via joint attention, improving closed‑loop performance on the nuPlan benchmark, and a training‑free guidance scheme called Decoupled Annealing Posterior Sampling with Energy (DAPSE) that injects arbitrary energy functions at inference time. Using the same model, the authors generate realistic long‑tail driving interactions—such as aggressive cut‑ins and lead‑vehicle braking—through inference‑time guidance, exposing failure modes in black‑box planners that standard benchmarks miss.
By Arka Pal, Rajesh Kumar, Hannes Eriksson, R\'emi Lacombe, Arvid Laveno Ling, Ankit Gupta, Maciej Wozniak
arXiv:2609.25831v1 Announce Type: cross
Abstract: Recent VLM-based autonomous driving planners adopt GRPO-style reinforcement learning to optimize driving performance. However, existing GRPO recipes...
By Yuqi Ye, Shangkun Sun, Junhong Lin, Jiayi Zhao, Changhao Peng, Wei Zheng, Guoqing Liu, Tiesong Zhao, Wei Gao