arXiv:2608.24167v1 Announce Type: new
Abstract: Whittle index policies offer a scalable method for restless multi-armed bandits, but under partial observability even determining the indifference subs...
By Qizhen Jia, Keqin Liu
arXiv:2409. 05980v2 Announce Type: replace-cross Abstract: Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolves over time due to the actions we perform or due to the nature.
By Gianmarco Genalti, Marco Mussi, Nicola Gatti, Marcello Restelli, Matteo Castiglioni, Alberto Maria Metelli
arXiv:2512. 09850v2 Announce Type: replace Abstract: We introduce Conformal Bandits, a novel framework integrating Conformal Prediction (CP) into bandit problems, a classic paradigm for sequential decision-making under uncertainty.
By Simone Cuonzo, Nina Deliu
arXiv:2602. 06014v2 Announce Type: replace-cross Abstract: Thompson sampling (TS) is widely used for stochastic multi-armed bandits, yet its inferential properties under adaptive data collection are subtle.
By Shunxing Yan, Han Zhong
arXiv:2307. 03587v4 Announce Type: replace Abstract: In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator.
By Nicklas Werge, Yi-Shan Wu, Abdullah Akg\"ul, Melih Kandemir
The paper introduces a new algorithm for a nonstationary bandit setting where actions influence both immediate rewards and the evolution of an unobserved latent linear state. By approximating the infinite‑memory reward process with a finite‑memory block‑level proxy and applying a UCB‑based block algorithm, the authors achieve a regret bound of “~O(√T)”, improving upon the previous “~O(T^{2/3})” guarantee. This represents the first such “~O(√T)” result for latent linear‑dynamics bandits with bilinear rewards and an open‑loop action‑sequence benchmark.
By Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh