Auto-exploration for online reinforcement learning
arXiv:2512. 06244v2 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms.
arXiv:2606. 15247v1 Announce Type: cross Abstract: The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting.
arXiv:2512. 06244v2 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms.
The article argues that Monte Carlo Tree Search (MCTS) and every‑visit Monte Carlo (MC) control are essentially the same method, differing only in terminology and presentation. It shows that MCTS’s four stages—selection, expansion, simulation, and backup—can be reduced to two core operations: sampling trajectories under the current policy and performing every‑visit MC updates. The note aims to make this equivalence explicit and easier to recognize.
arXiv:2608. 09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time.
arXiv:2609.39837v1 Announce Type: new Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
arXiv:2603. 16842v2 Announce Type: replace Abstract: Stochastic resetting -- intermittently returning a process to a fixed reference state -- has emerged as an effective mechanism for optimizing first-passage properties.
Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.
arXiv:2606. 10580v1 Announce Type: cross Abstract: The asymptotic behaviour of Monte Carlo optimistic policy iteration (MC-O-PI) is a long-standing open question.
arXiv:2602. 09300v2 Announce Type: replace Abstract: We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective.
arXiv:2602. 00781v2 Announce Type: replace Abstract: Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time.
arXiv:2607. 21637v1 Announce Type: new Abstract: This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments.
arXiv:2607. 17823v1 Announce Type: new Abstract: Reinforcement Learning is a cornerstone technique for modern large reasoning models.
arXiv:2607. 17201v1 Announce Type: cross Abstract: In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning.