arXiv Machine Learning By Volodymyr Tkachuk, Csaba Szepesv\'ari, Xiaoqi Tan

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability

Read the original on arXiv Machine Learning →

arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.