arXiv:2607. 17316v1 Announce Type: new Abstract: The softmax policy $\pi(a \mid s) \propto \exp(\beta Q(s,a))$ is the default model of stochastic choice in reinforcement learning (RL).
By Silviu Pitis
The paper introduces a Bellman optimality equation for optimizing plasticity in Markov decision processes, extending recent work that reframes the stability‑plasticity dilemma as an empowerment‑plasticity tradeoff. It builds on Abel et al. (2025), which defined plasticity as the generalized directed information from observations to actions and empowerment as the reverse. This work is the first to address plasticity optimization under that new definition, providing a theoretical foundation similar to existing empowerment‑based approaches.
By Jeremy Lucas, Doina Precup
arXiv:2602. 03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events.
By Aneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon, Erick Delage
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first st...
arXiv:2609.24103v1 Announce Type: new
Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. W...
By Larry Preuett, Qiuyi Zhang, Muhammad Aurangzeb Ahmad