arXiv Machine Learning

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

The paper introduces Odds‑Ratio Thompson Sampling (OR‑TS), a method for batched multi‑armed bandits that updates the joint posterior over log‑odds contrasts and refits the shared level in each batch, rather than carrying over absolute reward rates. It presents a Bayesian bandit agent with controls for decay of past evidence and aggressiveness of allocation, and evaluates OR‑TS against traditional absolute‑rate memory across 86 public A/B series and synthetic environments. Results show that when the shared level varies significantly, OR‑TS outperforms absolute‑rate memory, reducing regret and ensuring the best arm receives more traffic, while also handling cases where contrasts themselves shift.

arXiv AI
3d ago

Reserve-Aware Contrast Certificates for Conservative Bandits with Uncertain Baselines

The paper introduces Reserve-C4B, a method for conservative bandits that ensures improvement over an incumbent policy while respecting a performance budget, even when the incumbent’s reward is uncertain. By focusing on the baseline-relative contrast and using a shared confidence set, the approach derives an exact expression for the avoidable penalty and a tighter admissibility test at each history. The method incorporates a reserve ledger to separate statistical evidence from performance deficit and a prefix-refresh extension to recertify decisions without discarding prior credit, achieving high-probability conditional-mean performance guarantees for linear rewards.

By Qinchuan Cheng
arXiv Machine Learning
Sep 16

Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits

The paper introduces Decision‑Relevant Fresh Comparison (DRFC), a method for decentralized bandit systems with heterogeneous agents whose local reward changes may not affect the global best action. DRFC gathers balanced samples from all agents and only switches the common best arm when fresh global evidence indicates a change, yielding a dynamic regret bound that does not depend on the number of local changes. An anytime‑valid sliding‑window extension further handles gradual drift, and experiments on synthetic, semi‑real, and MovieLens‑1M data demonstrate that DRFC ignores decision‑irrelevant local changes while the extension avoids false switches.

By Zhaojun Peng
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv AI
Jun 9

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

arXiv:2606. 09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time.

By Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard