arXiv Machine Learning By Kaifei Wang, Yinyu Ye, Han Zhong

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

Read the original on arXiv Machine Learning →

arXiv:2608. 17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.