arXiv Machine Learning By Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Yougang Bian, Yongfu Li, Qisong Yang, Helai Huang

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

Read the original on arXiv Machine Learning →

arXiv:2606. 26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 23

Safety-Regulated Transfer Reinforcement Learning with Adaptive Teacher Guidance

arXiv:2606. 26527v2 Announce Type: replace Abstract: We propose Safety-Regulated Adaptive Transfer Reinforcement Learning (SRATRL), a teacher--student framework that combines safety-triggered intervention, safety-adaptive value shaping, and policy-compatibility-based optimization for efficient target-domain adaptation.

By Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Zeyu Yang, Qisong Yang, Yougang Bian
arXiv AI
Aug 12

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

arXiv:2608. 10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes.

By Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University)
arXiv Machine Learning
Sep 18

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

OPTED is a method for on‑policy fine‑tuning of end‑to‑end driving models that separates reinforcement learning from the policy update. A privileged teacher trained with RL on vectorized inputs (HD‑maps and bounding boxes) supervises the pre‑trained student during closed‑loop post‑training. Applied to the camera‑based models TransFuser and VaVAM in AlpaSim, OPTED boosts driving scores by 1.6× and 9.5×, respectively, while requiring roughly three orders of magnitude fewer simulator interactions than direct RL post‑training.

By Damiano Da Col, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, Christos Sakaridis