arXiv:2608.24167v1 Announce Type: new
Abstract: Whittle index policies offer a scalable method for restless multi-armed bandits, but under partial observability even determining the indifference subs...
By Qizhen Jia, Keqin Liu
arXiv:2409. 05980v2 Announce Type: replace-cross Abstract: Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolves over time due to the actions we perform or due to the nature.
By Gianmarco Genalti, Marco Mussi, Nicola Gatti, Marcello Restelli, Matteo Castiglioni, Alberto Maria Metelli
arXiv:2512. 09850v2 Announce Type: replace Abstract: We introduce Conformal Bandits, a novel framework integrating Conformal Prediction (CP) into bandit problems, a classic paradigm for sequential decision-making under uncertainty.
By Simone Cuonzo, Nina Deliu
arXiv:2602. 06014v2 Announce Type: replace-cross Abstract: Thompson sampling (TS) is widely used for stochastic multi-armed bandits, yet its inferential properties under adaptive data collection are subtle.
By Shunxing Yan, Han Zhong
arXiv:2307. 03587v4 Announce Type: replace Abstract: In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator.
By Nicklas Werge, Yi-Shan Wu, Abdullah Akg\"ul, Melih Kandemir
The paper introduces a new algorithm for a nonstationary bandit setting where actions influence both immediate rewards and the evolution of an unobserved latent linear state. By approximating the infinite‑memory reward process with a finite‑memory block‑level proxy and applying a UCB‑based block algorithm, the authors achieve a regret bound of “~O(√T)”, improving upon the previous “~O(T^{2/3})” guarantee. This represents the first such “~O(√T)” result for latent linear‑dynamics bandits with bilinear rewards and an open‑loop action‑sequence benchmark.
By Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh
arXiv:2604. 08149v2 Announce Type: replace Abstract: We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain.
By Zhen Li (LMO, CELESTE, HEC Paris), Gilles Stoltz (LMO, CELESTE, HEC Paris)
arXiv:2609. 22690v1 Announce Type: new Abstract: We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems.
By Huikang Liu, Zhengchao Wang, Daniel Kuhn, Wolfram Wiesemann
We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time. Under practitioner-friendly assumptions, we reduce this setting to linear bandit with stationary mean but heteroskedastic and non-stationary noise.
arXiv:2606. 09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time.
By Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard
arXiv:2511.08097v2 Announce Type: replace-cross
Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...
By Dheeraj Narasimha, Nicolas Gast
arXiv:2609.38132v1 Announce Type: new
Abstract: We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple...
By Yige Hong, Xiangcheng Zhang, Qiaomin Xie, Yudong Chen, Weina Wang