arXiv Machine Learning

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

arXiv:2608. 17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs.

arXiv Machine Learning
Jul 9

Nonlinear Bandit

arXiv:2607. 07304v1 Announce Type: new Abstract: In this paper we first study the problem of generalized linear bandit (GLB) under heavy-tailed noise.

By Tianshuo Zheng, Ting Wu, Zhi-Hua Zhou, Keqin Liu