arXiv Machine Learning By Weikai Wang, Erick Delage

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

Read the original on arXiv Machine Learning →

arXiv:2607. 23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.