Hugging Face Trending Papers

Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

Read the original on Hugging Face Trending Papers →

Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentially different transition probabilities and rewards.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.