arXiv Machine Learning

Safety-Regulated Transfer Reinforcement Learning with Adaptive Teacher Guidance

arXiv:2606. 26527v2 Announce Type: replace Abstract: We propose Safety-Regulated Adaptive Transfer Reinforcement Learning (SRATRL), a teacher--student framework that combines safety-triggered intervention, safety-adaptive value shaping, and policy-compatibility-based optimization for efficient target-domain adaptation.

arXiv Machine Learning
Jun 26

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

arXiv:2606. 26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making.

By Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Yougang Bian, Yongfu Li, Qisong Yang, Helai Huang
arXiv Machine Learning
Sep 21

Provably Optimal Reinforcement Learning under Safety Filtering

The paper proves that using a permissive safety filter in reinforcement learning does not compromise asymptotic performance. By formalizing safety through a safety‑critical Markov decision process and a filtered MDP, the authors show that optimal policies in the filtered MDP achieve the same return as the best safe policy in the original setting. Experiments on Safety Gymnasium confirm zero violations during training and performance that matches or exceeds unfiltered baselines.

By Donggeon David Oh, Duy P. Nguyen, Haimin Hu, Jaime Fern\'andez Fisac
arXiv Machine Learning
Jul 27

Safe In-Context Reinforcement Learning

arXiv:2509. 25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history.

By Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
arXiv AI
Sep 25

Beyond Average Safety: Chance-Constrained LLM Fine-tuning

The paper introduces a chance-constrained approach to fine‑tune large language models (LLMs) that limits the proportion of safety examples whose performance degrades beyond a set threshold relative to a reference model. By replacing the discontinuous violation indicator with a differentiable majorization, the authors derive a tractable, conservative constraint and a closed‑form, constraint‑aware gradient update that focuses on examples near or above the degradation threshold. Experiments on harmful fine‑tuning across three tasks and models show that this tail‑aware method consistently outperforms existing safety‑preserving baselines, suggesting that safety preservation should be treated as a reliability‑constrained optimization problem rather than average‑risk regularization.

By Taha Entesari, Mahyar Fazlyab