arXiv Machine Learning By Abdelghani Ghanem, Mounir Ghogho

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 02034v1 Announce Type: new Abstract: Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, inducing a pessimistic bias that grows with the horizon.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.