arXiv:2606. 09002v1 Announce Type: cross Abstract: We study a stochastic multi-armed bandit problem in which the set of available arms expands over time.
By Deqi Zheng, Xiaoyang Xu, Yuhong Yang
arXiv:2604. 08149v2 Announce Type: replace Abstract: We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain.
By Zhen Li (LMO, CELESTE, HEC Paris), Gilles Stoltz (LMO, CELESTE, HEC Paris)
arXiv:2501. 19401v5 Announce Type: replace Abstract: We introduce a practical, black-box framework termed Detection Augmented Learning (DAL) for the problem of piecewise stationary bandits without knowledge of the underlying non-stationarity.
By Argyrios Gerogiannis, Yu-Han Huang, Subhonmesh Bose, Venugopal V. Veeravalli
arXiv:2606. 23933v1 Announce Type: cross Abstract: We study non-stationary linear contextual bandits where the reward model drifts over time, rendering classical contextual bandit algorithms brittle because historical data becomes systematically biased.
By AmirHossein Naghdi, Ali Baheri
The paper extends the idea that contexts are cheap for linear bandits from i.i.d. settings to Markovian context processes. By assuming uniform geometric ergodicity, the authors construct a stationary surrogate action set and use a delayed‑update scheme to mitigate bias from nonstationary conditional context distributions. They provide a phased algorithm for unknown stationary distributions and achieve high‑probability regret bounds comparable to standard linear bandit oracles in fast‑mixing regimes, with empirical validation showing gains over LinUCB.
By Kaan Buyukkalayci, Osama Hanna, Christina Fragouli
The paper introduces Odds‑Ratio Thompson Sampling (OR‑TS), a method for batched multi‑armed bandits that updates the joint posterior over log‑odds contrasts and refits the shared level in each batch, rather than carrying over absolute reward rates. It presents a Bayesian bandit agent with controls for decay of past evidence and aggressiveness of allocation, and evaluates OR‑TS against traditional absolute‑rate memory across 86 public A/B series and synthetic environments. Results show that when the shared level varies significantly, OR‑TS outperforms absolute‑rate memory, reducing regret and ensuring the best arm receives more traffic, while also handling cases where contrasts themselves shift.
By Sulgi Kim