arXiv AI

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit

arXiv:2609. 38659v1 Announce Type: cross Abstract: We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers.