UCB exploration via Q-ensembles
Related stories
An Introduction to Q-Learning Part 2/2
Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning
arXiv:2606. 20107v1 Announce Type: new Abstract: Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration.
Information-Based Exploration via Random Features for Reinforcement Learning
arXiv:2607. 17981v1 Announce Type: new Abstract: Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical guarantees harder to establish.
Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments
arXiv:2608.21995v1 Announce Type: cross Abstract: We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exploration p...
Constraint-Enhanced Physical Search through Correlation Matching
arXiv:2606. 03554v1 Announce Type: cross Abstract: Physical systems do not merely add noise to search processes; they impose constraints that generate structured correlations.
#Exploration: A study of count-based exploration for deep reinforcement learning
Bandits via Additive Quantized Representations
The paper introduces Residual Quantization (RQ) as a representation layer for contextual bandits, converting continuous contexts into discrete centroid assignments across multiple levels. This approach allows additive bandit algorithms to capture nonlinear reward structures while maintaining strictly bounded memory and efficient online updates. Experiments on 13 datasets show that RQ variants outperform non-RQ counterparts on 11 datasets and match advanced baselines like XGBoost and neural methods using up to 1000 times less memory.
Mixture-Greedy for Online Generative Model Selection: Is UCB Necessary in Diversity-Aware Multi-Armed Bandits?
arXiv:2603.21716v2 Announce Type: replace-cross Abstract: Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models...
Deep Q-Learning with Space Invaders
GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling
arXiv:2609.14711v2 Announce Type: replace Abstract: Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport pro...
Randomized Exploration for Linear Bandits via Absolute Perturbations
arXiv:2606. 28616v1 Announce Type: new Abstract: In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson Sampling (TS) is computationally attractive yet typically harder to analyze due to its non-optimistic nature.