arXiv Machine Learning By Jingxin Zhan, Yuze Han, Zhihua Zhang

Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?

Read the original on arXiv Machine Learning →

arXiv:2608. 15365v1 Announce Type: new Abstract: Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
5d ago

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity

The paper investigates multi‑armed bandits where several arms are optimal. It refines previous sub‑sampling algorithms to achieve a minimax regret of τO((K−A)/√(KA)·√T), improving on earlier bounds. A matching lower bound is provided, showing the rate is nearly optimal, and the authors demonstrate that knowing the number of optimal arms A (within constant factors) is essential for near‑optimal performance.

By Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu