arXiv Machine Learning By Larissa Xu, King Bi, William Chang

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

Read the original on arXiv Machine Learning →

arXiv:2608. 12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.