In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-value) functions are treated as uncertain quantities, and their estimation is formulated as a stochastic inference problem.
arXiv:2607. 20010v1 Announce Type: new Abstract: In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations.
By Vasos Arnaoutis, Eric Lutters, Bojana Rosi\'c
arXiv:2606. 05967v1 Announce Type: cross Abstract: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA).
By Ziad Kobeissi (L2S), \'Elo\"ise Berthier (U2IS)
arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.
By Mohsen Amiri, Sindri Magn\'usson
arXiv:2609.06882v1 Announce Type: cross
Abstract: Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remain...
By Mahmoud Selim, Cristina Cipriani, Karl H. Johansson
arXiv:2505.01361v3 Announce Type: replace
Abstract: Temporal difference (TD) learning is a foundational algorithm in reinforcement learning (RL). For nearly forty years, TD learning has served as a w...
By Hwanwoo Kim, Panos Toulis, Eric Laber
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
arXiv:2607. 22399v1 Announce Type: cross Abstract: We consider the problem of learning from a single finite trajectory of an ergodic stochastic dynamical system.
By Oleksii Kachaiev, Silvia Villa, Lorenzo Rosasco
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
The paper investigates the finite‑iteration behavior of exact asynchronous recursions used in categorical distributional temporal‑difference (TD) learning. It analyzes both scalar categorical TD in the Cramér geometry and multivariate signed‑categorical TD in the maximum mean discrepancy geometry, showing that these methods can be viewed as single‑state stochastic‑approximation recursions that contract in a block‑supremum norm. The authors develop a restricted‑domain theory, derive discounted bounds under i.i.d. and Markovian sampling, and extend the analysis to undiscounted fixed‑horizon policy evaluation with horizon‑stacked categorical methods under episodic sampling, thereby providing a unified non‑asymptotic analysis across various settings.
By Ege C. Kaya, Abolfazl Hashemi
arXiv:2606. 16846v1 Announce Type: cross Abstract: We study the operator-theoretic core of Q-learning in continuous-time stochastic control with continuous states and actions.
By Qian Qi
arXiv:2607. 07967v1 Announce Type: cross Abstract: Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, representing state-conditioned action distributions as terminal laws of diffusion processes with parameterized drifts.
By Viet Vu, Renyuan Xu, Jiacheng Zhang, Yufei Zhang