The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.
By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
arXiv:2606. 08369v1 Announce Type: cross Abstract: A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment.
By Wanqiao Xu, Yifan Zhu, Benjamin Van Roy
arXiv:2608. 10529v1 Announce Type: cross Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions.
By Daphne Feng, Ricardo Parada, Lily Jiang, Sophia Yi, William Chang
arXiv:2602. 12963v2 Announce Type: replace Abstract: An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world.
By Alfred Harwood, Jose Faustino, Alex Altair
arXiv:2607. 29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process.
By Bumgeun Park, Donghwan Lee
The paper investigates when information sharing enhances decentralized discovery by separating its effects on pooled estimation and independent rescue actions in finite discovery models. It shows that a registered incremental-sharing protocol improves discovery only when pooled residual error decreases faster than an independent rescue attempt, and that equilibrium selection can determine whether sharing is beneficial. The study uses synthetic, finite models without human or organizational data.
By Yohei Nakajima