arXiv Machine Learning By Timothy Tomashevskiy

Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2607. 21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 26

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

arXiv:2606. 26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making.

By Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Yougang Bian, Yongfu Li, Qisong Yang, Helai Huang
arXiv Machine Learning
Jul 23

Safety-Regulated Transfer Reinforcement Learning with Adaptive Teacher Guidance

arXiv:2606. 26527v2 Announce Type: replace Abstract: We propose Safety-Regulated Adaptive Transfer Reinforcement Learning (SRATRL), a teacher--student framework that combines safety-triggered intervention, safety-adaptive value shaping, and policy-compatibility-based optimization for efficient target-domain adaptation.

By Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Zeyu Yang, Qisong Yang, Yougang Bian
arXiv Machine Learning
Jul 27

Safe In-Context Reinforcement Learning

arXiv:2509. 25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history.

By Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
arXiv Machine Learning
Sep 21

Provably Optimal Reinforcement Learning under Safety Filtering

The paper proves that using a permissive safety filter in reinforcement learning does not compromise asymptotic performance. By formalizing safety through a safety‑critical Markov decision process and a filtered MDP, the authors show that optimal policies in the filtered MDP achieve the same return as the best safe policy in the original setting. Experiments on Safety Gymnasium confirm zero violations during training and performance that matches or exceeds unfiltered baselines.

By Donggeon David Oh, Duy P. Nguyen, Haimin Hu, Jaime Fern\'andez Fisac