arXiv Machine Learning By Toshinori Kitamura, Shuai Liu, Csaba Szepesv\'ari

Randomized Exploration for Linear Bandits via Absolute Perturbations

Read the original on arXiv Machine Learning →

arXiv:2606. 28616v1 Announce Type: new Abstract: In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson Sampling (TS) is computationally attractive yet typically harder to analyze due to its non-optimistic nature.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.