arXiv Machine Learning

Adaptive Exploration for Latent-State Bandits

arXiv:2602. 05139v3 Announce Type: replace Abstract: We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions.

arXiv Machine Learning
Sep 10

Mixing Makes Markovian Contexts Cheap for Linear Bandits

The paper extends the idea that contexts are cheap for linear bandits from i.i.d. settings to Markovian context processes. By assuming uniform geometric ergodicity, the authors construct a stationary surrogate action set and use a delayed‑update scheme to mitigate bias from nonstationary conditional context distributions. They provide a phased algorithm for unknown stationary distributions and achieve high‑probability regret bounds comparable to standard linear bandit oracles in fast‑mixing regimes, with empirical validation showing gains over LinUCB.

By Kaan Buyukkalayci, Osama Hanna, Christina Fragouli
arXiv Machine Learning
Sep 18

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

The paper introduces Odds‑Ratio Thompson Sampling (OR‑TS), a method for batched multi‑armed bandits that updates the joint posterior over log‑odds contrasts and refits the shared level in each batch, rather than carrying over absolute reward rates. It presents a Bayesian bandit agent with controls for decay of past evidence and aggressiveness of allocation, and evaluates OR‑TS against traditional absolute‑rate memory across 86 public A/B series and synthetic environments. Results show that when the shared level varies significantly, OR‑TS outperforms absolute‑rate memory, reducing regret and ensuring the best arm receives more traffic, while also handling cases where contrasts themselves shift.

By Sulgi Kim
arXiv AI
Jun 9

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

arXiv:2606. 09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time.

By Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard