arXiv AI By Yang Xu, Swetha Ganesh, Vaneet Aggarwal

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

Read the original on arXiv AI →

arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.