arXiv Machine Learning By Sterre Lutz, Dani\"el Vos, Matthijs T. J. Spaan, Anna Lukina

Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

Read the original on arXiv Machine Learning →

arXiv:2608. 02509v1 Announce Type: cross Abstract: Sequential decision-making in real-world applications often involves uncertainty about the environment's model.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.