Learning to Plan from Random Exploration
arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without po...
The paper offers a new way to view the occupancy measure in reinforcement learning by embedding the planning criterion into the dynamics via a resetting planning process. The resulting stationary measure, called the visitation measure, forms a dually flat statistical manifold with two affine charts: visitation probabilities and log-policies, which are dual under conditional entropy. This geometric framework allows planning-as-inference to extend beyond linear rewards to nonlinear functionals of visitation, with each iteration solvable by a natural-gradient step and provides a new interpretation of the temporal-difference error as a marginal-utility estimate.
arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without po...
The paper develops a geometric theory of decision boundaries for structured Markov Decision Processes, treating the geometry induced by optimal policies as the key analytical object. It shows that, under structural regularity, this geometry yields the minimal representation needed for policy reconstruction and dictates the statistical and computational complexity of the reconstruction problem. The authors introduce intrinsic notions of boundary and decision complexity, derive information-theoretic measures of decision compression, and provide statistical guarantees for boundary estimation and policy reconstruction from black-box queries, supported by controlled numerical experiments.
arXiv:2602. 17315v3 Announce Type: replace-cross Abstract: We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, where accessibility of the next action is restricted to a subset dependent on the agent's current choice.
The paper discusses how reinforcement learning theory relies on probability theory via Markov chains and highlights a deep link between probability theory and potential theory. It reviews this connection and examines how a potential-theoretic perspective can be applied to core RL representations and algorithms under a fixed‑policy assumption, suggesting possible gains in sample efficiency and formal constraints. The authors also note that relaxing the fixed‑policy assumption allows the linear potential theory framework to extend naturally to nonlinear cases.
arXiv:2510. 02149v2 Announce Type: replace Abstract: We introduce Action-Triggered Sporadically Traceable Markov Decision Processes (ATST-MDPs), a reinforcement learning framework for partial observability in which full state observations occur stochastically at each step, with probability determined by the chosen action.
The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.
arXiv:2610.01413v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and e...
The paper introduces ReQRL, a method for learning quasimetric geometry in goal-conditioned reinforcement learning by constraining the critic’s value gradients with finite-horizon reachability. It decouples dynamical reachability from boundary geometry, estimating both from data using state-constrained optimal control principles. Experiments on OGBench show that ReQRL matches or surpasses existing quasimetric and offline GCRL approaches.
arXiv:2602. 05031v2 Announce Type: replace Abstract: Planning with a learned model remains a key challenge in model-based reinforcement learning (RL).
arXiv:2505.04193v2 Announce Type: replace Abstract: Simplicity is a critical inductive bias for designing data-driven controllers, especially when robustness is important. Despite the impressive resu...
arXiv:2606. 18531v1 Announce Type: cross Abstract: Offline reinforcement learning is typically analyzed under process-level reward supervision, yet many sequential decision datasets record only trajectory-level outcomes.
arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.