arXiv AI By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang

Robust Regularized Policy Iteration under Transition Uncertainty

Read the original on arXiv AI →

arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.