arXiv Machine Learning By Mark Sellke, Gregory Valiant

Thompson Sampling Is 2-Competitive for Mistakes

Read the original on arXiv Machine Learning →

arXiv:2607. 12389v1 Announce Type: cross Abstract: We consider Bayesian bandit models and prove that Thompson sampling makes at most twice the expected number of mistakes (selections of a suboptimal arm) as any other policy.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.