arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.
By Soichiro Nishimori, Paavo Parmas
REVERSAL-BENCH is a benchmark that introduces a continuous reversibility parameter ρ∈[0,1] and a reset oracle to evaluate how well reinforcement learning agents can recover from irreversible states across eight manipulation tasks in five physics engines. Experiments show a sharp reversibility cliff: reset‑free agents become trapped in irrecoverable states as ρ increases, while episodic agents continue learning steadily. The benchmark also provides a large multi‑simulator dataset and demonstrates that safety shields can predict recoverability but only succeed when the agent can avoid the trap.
By Riyaaz Shaik, Chandru Venkataraman
arXiv:2606. 15247v1 Announce Type: cross Abstract: The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting.
By Octave Oliviers, Glenn Vinnicombe
arXiv:2512.08463v2 Announce Type: replace
Abstract: We study how privileged information about a physical system affects the discovery of high-performing policies when training a reinforcement learnin...
By Antonio Terpin, Raffaello D'Andrea
arXiv:2505. 18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fixed for evaluation.
By Mohamed A. Mohamed, Kateryna Nekhomiazh, Vedant Vyas, Marcos M. Jose, Andrew Patterson, Marlos C. Machado
arXiv:2606. 08452v1 Announce Type: new Abstract: In many real-world settings, data streams are nonstationary and arrive sequentially, requiring learning systems to adapt continuously without retraining from scratch.
By Nazreen Shah, Govinda Arya, Bharath B. N., Ranjitha Prasad