OpenAI Blog

UCB exploration via Q-ensembles

arXiv Machine Learning
1d ago

Bandits via Additive Quantized Representations

The paper introduces Residual Quantization (RQ) as a representation layer for contextual bandits, converting continuous contexts into discrete centroid assignments across multiple levels. This approach allows additive bandit algorithms to capture nonlinear reward structures while maintaining strictly bounded memory and efficient online updates. Experiments on 13 datasets show that RQ variants outperform non-RQ counterparts on 11 datasets and match advanced baselines like XGBoost and neural methods using up to 1000 times less memory.

By Ami Tavory, Noam Touitou, Tal Sarig, Frank Cheng, Ido Guy
arXiv Machine Learning
Jun 30

Randomized Exploration for Linear Bandits via Absolute Perturbations

arXiv:2606. 28616v1 Announce Type: new Abstract: In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson Sampling (TS) is computationally attractive yet typically harder to analyze due to its non-optimistic nature.

By Toshinori Kitamura, Shuai Liu, Csaba Szepesv\'ari