Human-like autonomy emerges from self-play and a pinch of human data
arXiv:2606. 19370v1 Announce Type: cross Abstract: Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data.
The paper reports on training autonomous driving policies via self‑play, extending previous work by using Transformers and a real‑city high‑definition map. On CARLA and Waymo benchmarks, the resulting policies underperform compared to Gigaflow, with identified failure modes such as reward hacking at traffic lights and lack of incentive to stop at stop signs. The authors also analyze which traffic rules emerge from self‑play and confirm that reward conditioning produces diverse driving behaviors.
arXiv:2606. 19370v1 Announce Type: cross Abstract: Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data.
arXiv:2607. 13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains.
arXiv:2506. 12283v2 Announce Type: replace Abstract: Modeling vehicle interactions at unsignalized intersections is a challenging task due to the complexity of the underlying game-theoretic processes.
The paper introduces a plug‑and‑play method that injects traffic‑element signals—such as traffic lights and road signs—into end‑to‑end autonomous driving models with minimal architectural changes. By augmenting several public datasets with comprehensive traffic‑element annotations, the authors evaluate this integration across diverse driving paradigms, consistently improving performance on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. The approach achieves a new state‑of‑the‑art result on the challenging NAVSIM‑v2 benchmark, demonstrating the broad utility of traffic‑element awareness.
arXiv:2510. 12560v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving models trained with imitation learning (IL) often generalize poorly, particularly in long-tail scenarios where expert demonstrations are sparse.
arXiv:2608. 14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them.
arXiv:2606. 17362v1 Announce Type: cross Abstract: Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent.
arXiv:2608. 11451v1 Announce Type: cross Abstract: Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss.
End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in autonomous driving. This class of models includes both Vision-Language-Action models and trajectory-generative planners.
arXiv:2606. 31106v1 Announce Type: cross Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics.
arXiv:2412.02520v4 Announce Type: replace-cross Abstract: Connected automated vehicles (CAVs) equipped with adaptive cruise control (ACC) create new opportunities for highway congestion mitigation. T...
arXiv:2608. 06445v1 Announce Type: cross Abstract: Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation.