arXiv Machine Learning

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

arXiv:2607. 23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning.

arXiv Machine Learning
Jun 24

A Robust Model-Based Approach for Continuous-Time Policy Evaluation with Unknown L\'evy Process Dynamics

arXiv:2504. 01482v3 Announce Type: replace-cross Abstract: This paper develops a model-based framework for continuous-time policy evaluation (CTPE) in reinforcement learning, incorporating both Brownian and L\'evy noise to model stochastic dynamics influenced by rare and extreme events.

By Qihao Ye, Xiaochuan Tian, Yuhua Zhu