arXiv Machine Learning By Pu Wang, Yao-Xiang Ding

Tree-Guided Identify-Then-Exploit: A Unified Framework of Best Arm Identification and Regret Minimization for Dueling Bandits

Read the original on arXiv Machine Learning →

arXiv:2606. 01799v1 Announce Type: new Abstract: We study $N$-armed stochastic dueling bandits under the Condorcet-winner assumption, where three widely adopted objectives are considered: best-arm identification (BAI), weak regret, and strong regret.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.