arXiv Machine Learning By Sreejeet Maity, Aritra Mitra

Robust Asynchronous Q-Learning under Reward and State Corruption via Batching

Read the original on arXiv Machine Learning →

arXiv:2607. 20822v1 Announce Type: new Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.