arXiv Machine Learning

Restless bandits with imperfect binary feedback: PCL-indexability analysis and computation

arXiv:2606. 11192v1 Announce Type: new Abstract: We study restless bandits with binary latent states and imperfect binary feedback, motivated by opportunistic spectrum access with sensing errors.

arXiv Machine Learning
Jun 30

Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting

arXiv:2409. 05980v2 Announce Type: replace-cross Abstract: Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolves over time due to the actions we perform or due to the nature.

By Gianmarco Genalti, Marco Mussi, Nicola Gatti, Marcello Restelli, Matteo Castiglioni, Alberto Maria Metelli
arXiv Statistics ML
4d ago

Block Optimism for Nonstationary Bandits with Latent Linear Dynamics

The paper introduces a new algorithm for a nonstationary bandit setting where actions influence both immediate rewards and the evolution of an unobserved latent linear state. By approximating the infinite‑memory reward process with a finite‑memory block‑level proxy and applying a UCB‑based block algorithm, the authors achieve a regret bound of “~O(√T)”, improving upon the previous “~O(T^{2/3})” guarantee. This represents the first such “~O(√T)” result for latent linear‑dynamics bandits with bilinear rewards and an open‑loop action‑sequence benchmark.

By Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh
Hugging Face Trending Papers
Jun 8

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time. Under practitioner-friendly assumptions, we reduce this setting to linear bandit with stationary mean but heteroskedastic and non-stationary noise.

arXiv AI
Jun 9

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

arXiv:2606. 09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time.

By Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard