Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentially different transition probabilities and rewards.
arXiv:2608. 02509v1 Announce Type: cross Abstract: Sequential decision-making in real-world applications often involves uncertainty about the environment's model.
By Sterre Lutz, Dani\"el Vos, Matthijs T. J. Spaan, Anna Lukina
Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment.
arXiv:2608. 17929v1 Announce Type: new Abstract: Robust Markov decision processes optimize one policy against a set of plausible transition functions.
By Kasper Engelen, Sebastian Junges, Guillermo A. P\'{e}rez, Marnix Suilen
arXiv:2606. 14095v1 Announce Type: new Abstract: We study the sample complexity of learning in average-reward weakly-coupled Markov decision processes (WCMDPs) and Restless Bandits (RBs) under a generative model.
By Tianhao Wu, Matthew Zurek, Weina Wang, Qiaomin Xie
arXiv:2509. 16586v2 Announce Type: replace Abstract: Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model.
By Yukuan Wei, Xudong Li, Lin F. Yang