arXiv:2602. 17315v3 Announce Type: replace-cross Abstract: We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, where accessibility of the next action is restricted to a subset dependent on the agent's current choice.
By Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni, Lijun Chen
arXiv:2409. 05980v2 Announce Type: replace-cross Abstract: Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolves over time due to the actions we perform or due to the nature.
By Gianmarco Genalti, Marco Mussi, Nicola Gatti, Marcello Restelli, Matteo Castiglioni, Alberto Maria Metelli
arXiv:2606. 08977v1 Announce Type: new Abstract: Motivated by the recency effect in online learning, we study algorithms for single-pass *sliding-window streaming multi-armed bandits (MABs)* in this paper.
By Vladimir Braverman, Chen Wang, Liudeng Wang, Samson Zhou
arXiv:2606. 09002v1 Announce Type: cross Abstract: We study a stochastic multi-armed bandit problem in which the set of available arms expands over time.
By Deqi Zheng, Xiaoyang Xu, Yuhong Yang
The paper introduces Decision‑Relevant Fresh Comparison (DRFC), a method for decentralized bandit systems with heterogeneous agents whose local reward changes may not affect the global best action. DRFC gathers balanced samples from all agents and only switches the common best arm when fresh global evidence indicates a change, yielding a dynamic regret bound that does not depend on the number of local changes. An anytime‑valid sliding‑window extension further handles gradual drift, and experiments on synthetic, semi‑real, and MovieLens‑1M data demonstrate that DRFC ignores decision‑irrelevant local changes while the extension avoids false switches.
By Zhaojun Peng
Decentralized bandit systems often contain heterogeneous agents: rewards can change at individual agents even when the best action for the network stays the same. These local changes may cancel when r...
arXiv:2604. 00523v2 Announce Type: replace Abstract: We study for the first time, stochastic dueling bandits over continuous action spaces with Lipschitz structure, where feedback is purely comparative.
By Mudit Sharma, Shweta Jain, Vaneet Aggarwal, Ganesh Ghalme
We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal is to maximize the cumulative reward over time.
arXiv:2609.13547v1 Announce Type: new
Abstract: We study switching regret in adversarial multi-armed bandits, where the learner competes with an arm sequence that changes at most $S$ times. When $S$...
By Mengxiao Zhang
arXiv:2602.10727v3 Announce Type: replace
Abstract: Rising Multi-Armed Bandits (RMABs) model sequential decision problems where each arm's expected reward improves with repeated pulls. In such proble...
By Seockbean Song, Chenyu Gan, Youngsik Yoon, Siwei Wang, Wei Chen, Jungseul Ok
arXiv:2607. 13686v1 Announce Type: new Abstract: We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation.
By Hao Qin, Chicheng Zhang
arXiv:2606. 27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs.
By Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop