The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.
By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
arXiv:2609.08800v1 Announce Type: cross
Abstract: Three properties determine whether a differentiable simulator can drive gradient-based optimization through contact: simulation accuracy, gradient re...
By Ale\v{s} Ku\v{c}era, Karel Zimmermann
arXiv:2607. 13319v1 Announce Type: cross Abstract: High-speed off-road autonomy requires precise closed-loop control for a target vehicle while remaining robust across changing terrains.
By Rwik Rana, Jesse Quattrociocchi, Christian Ellis, Nathan Tsoi, Garrett Warnell, Joydeep Biswas
arXiv:2605. 04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning.
By Jonathan Spieler, Sven Behnke
The paper introduces Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that corrects intermediate waypoints of end-to-end driving policies while preserving the predicted endpoint. ECO does not require maps, privileged simulator state, or additional training, and can be applied to a wide range of waypoint-emitting policies. Experiments on two closed-loop simulators show that ECO significantly improves closed-loop performance, achieving top results in the HUGSIM Closed-Loop Driving Challenge and boosting scene scores on AlpaSim.
By Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta
arXiv:2609.13845v1 Announce Type: cross
Abstract: World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet pl...
By Saksham Bansal, Om Naphade, Chayan Aggarwal, Vrishin M
arXiv:2607. 13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains.
By Zhouchonghao Wu, Akshay Rangesh, Weixin Li, Wei-Jer Chang, Zachary Lee, Tim Wang, Wei Zhan
arXiv:2608. 02069v1 Announce Type: cross Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training times.
By Martin Opat
arXiv:2601. 00728v5 Announce Type: replace Abstract: We propose a reinforcement learning (RL) framework for \xy{responsive} precision tuning for linear solvers, which can be extended to general algorithms.
By Erin Carson, Xinye Chen
arXiv:2606.23079v2 Announce Type: replace-cross
Abstract: Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but t...
By Yutian Cheng, Xiaojian Ma, Xianhao Wang, Min Yang, Rongpeng Su, Hangxin Liu, Xi Chen, Shuai Li, Qing Li
SlackDrive is a pre‑inference compute allocator that dynamically selects the compute budget for each driving control step by reusing the realized latency from previous inferences. By profiling a small set of discrete budgets once, it estimates the current compute state online and chooses the highest‑utility budget that stays within the admissible latency envelope. On the NAVSIM v2 benchmark with DriveDreamer‑Policy, SlackDrive boosts latency‑constrained EPDMS performance by 21.7% compared to the best baseline, while full‑budget and token‑pruning approaches exceed the latency limits under runtime contention.
By Xiaohuan Pei, Hengguang Zhou, Yuanhao Ban, Justin Cui, Jiaqi Feng, Haoyu Xie, Tao Huang, Pichao Wang, Yanchao Yang, Cho-Jui Hsieh
arXiv:2606. 00383v1 Announce Type: cross Abstract: While Model Predictive Control (MPC) provides strong stability and robustness, it imposes a significant computational burden on real-time systems.
By Theo Guegan, Dexter Wen Jie Teo