← Back to all news
arXiv Machine Learning October 2, 2026 By Seockbean Song, Chenyu Gan, Youngsik Yoon, Siwei Wang, Wei Chen, Jungseul Ok

Rising Multi-Armed Bandits with Known Horizons

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 9

Multi-Armed Bandits with Arriving Arms: Sequential Screening, Dynamic Regret, and Sublinear Guarantees

arXiv:2606. 09002v1 Announce Type: cross Abstract: We study a stochastic multi-armed bandit problem in which the set of available arms expands over time.

By Deqi Zheng, Xiaoyang Xu, Yuhong Yang
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Sep 22

Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping

arXiv:2609. 22690v1 Announce Type: new Abstract: We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems.

By Huikang Liu, Zhengchao Wang, Daniel Kuhn, Wolfram Wiesemann
reinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jun 29

Learning in Markovian bandits with non-observable states and constrained decision epochs

arXiv:2606. 27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs.

By Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop
reinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 9

Online Learning with Recency: Algorithms for Sliding-window Streaming Multi-armed Bandits

arXiv:2606. 08977v1 Announce Type: new Abstract: Motivated by the recency effect in online learning, we study algorithms for single-pass *sliding-window streaming multi-armed bandits (MABs)* in this paper.

By Vladimir Braverman, Chen Wang, Liudeng Wang, Samson Zhou
More like this →
arXiv Machine Learning
Aug 4

Conformal bandits: bringing statistical validity and reward efficiency under weak arm separability

arXiv:2512. 09850v2 Announce Type: replace Abstract: We introduce Conformal Bandits, a novel framework integrating Conformal Prediction (CP) into bandit problems, a classic paradigm for sequential decision-making under uncertainty.

By Simone Cuonzo, Nina Deliu
reinforcement-learning
More like this →
arXiv Machine Learning
Jul 16

On the Sublinear Regret of Continuous K-Max Bandits

arXiv:2502. 13467v2 Announce Type: replace Abstract: The $K$-Max combinatorial multi-armed bandit problem arises in applications such as recommendation and distributed decision making, where the reward is determined by the maximum outcome among $K$ selected arms.

By Yu Chen, Siwei Wang, Longbo Huang, Wei Chen
reinforcement-learningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea