← Back to all news
arXiv Statistics ML October 2, 2026 By Nam Nguyen, Tuan Quang Dam

Sharp Non-Asymptotic Analysis of the Penalized Challenger in $\beta$-EB-TCI for Bernoulli Bandits

Read the original on arXiv Statistics ML →

The Flow has not summarised this story yet — read it at arXiv Statistics ML.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
3d ago

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit

arXiv:2609. 38659v1 Announce Type: cross Abstract: We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers.

By Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu
More like this →
arXiv AI
2d ago

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity

arXiv:2609.38659v2 Announce Type: replace-cross Abstract: We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multi...

By Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu
More like this →
arXiv Machine Learning
Jul 13

Optimal Top-$k$ Identification from Pairwise Comparisons

arXiv:2607. 08979v1 Announce Type: new Abstract: We study the active learning problem of fixed-confidence top-$k$ identification from noisy pairwise comparisons.

By Motti Goldberger, Nils Rudi
reinforcement-learning
More like this →
arXiv Machine Learning
Jul 7

Replicability is Asymptotically Free in Multi-armed Bandits

arXiv:2402. 07391v3 Announce Type: replace-cross Abstract: We consider a replicable stochastic multi-armed bandit algorithm that ensures, with high probability, that the algorithm's sequence of actions is not affected by the randomness inherent in the dataset.

By Junpei Komiyama, Shinji Ito, Yuichi Yoshida, Souta Koshino
reinforcement-learning
More like this →
arXiv Machine Learning
2d ago

Regret Analysis of Retry-Based Bandits

arXiv:2605.20854v3 Announce Type: replace Abstract: We provide the first regret analysis of ReMax in stochastic multi-armed bandits. Originally introduced for reinforcement learning, ReMax is motivat...

By Bingkui Tong, Junpei Komiyama, Soichiro Nishimori, Paavo Parmas
reinforcement-learning
More like this →
arXiv Machine Learning
Aug 18

Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits

arXiv:2510. 22819v3 Announce Type: replace Abstract: The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learner's actual decisions and describes the evolution of the learning process over time.

By Jingxin Zhan, Yuze Han, Zhihua Zhang
safety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea