In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-value) functions are treated as uncertain quantities, and their estimation is formulated as a stochastic inference problem.
arXiv:2607. 20010v1 Announce Type: new Abstract: In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations.
By Vasos Arnaoutis, Eric Lutters, Bojana Rosi\'c
arXiv:2606. 05967v1 Announce Type: cross Abstract: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA).
By Ziad Kobeissi (L2S), \'Elo\"ise Berthier (U2IS)
arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.
By Mohsen Amiri, Sindri Magn\'usson
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
arXiv:2607. 22399v1 Announce Type: cross Abstract: We consider the problem of learning from a single finite trajectory of an ergodic stochastic dynamical system.
By Oleksii Kachaiev, Silvia Villa, Lorenzo Rosasco
arXiv:2605. 06866v2 Announce Type: replace Abstract: We study finite-iteration behavior of the exact asynchronous recursions used by categorical distributional temporal-difference methods.
By Ege C. Kaya, Abolfazl Hashemi
arXiv:2606. 16846v1 Announce Type: cross Abstract: We study the operator-theoretic core of Q-learning in continuous-time stochastic control with continuous states and actions.
By Qian Qi
arXiv:2607. 07967v1 Announce Type: cross Abstract: Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, representing state-conditioned action distributions as terminal laws of diffusion processes with parameterized drifts.
By Viet Vu, Renyuan Xu, Jiacheng Zhang, Yufei Zhang
arXiv:2411. 01302v2 Announce Type: replace Abstract: We study the convergence of $q$-learning and related algorithms introduced by Jia and Zhou (J.
By Wenpin Tang, Xun Yu Zhou
arXiv:2606. 04275v1 Announce Type: cross Abstract: We present a novel theoretical framework for deep reinforcement learning (RL) in continuous environments by modeling the problem as a continuous-time stochastic process, drawing on insights from stochastic control.
By Saket Tiwari, Tejas Kotwal, George Konidaris
This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based formulation, policy evaluation is naturally described by a stationary Hamilton-Jacobi-Bellman equation on $\mathcal P_2(\mathbb R^d)$, but this equation involves the drift and diffusion coefficients of the controlled McKean-Vlasov dynamics, which are not identifiable when only discrete-time data are available.