arXiv AI

Generalised Bellman recurrence and three dualities in sequential decision-making

arXiv:2607. 18077v1 Announce Type: cross Abstract: What gives the Bellman equation its form?

arXiv Machine Learning
Sep 11

A Bellman Optimality Equation for Plasticity

The paper introduces a Bellman optimality equation for optimizing plasticity in Markov decision processes, extending recent work that reframes the stability‑plasticity dilemma as an empowerment‑plasticity tradeoff. It builds on Abel et al. (2025), which defined plasticity as the generalized directed information from observations to actions and empowerment as the reverse. This work is the first to address plasticity optimization under that new definition, providing a theoretical foundation similar to existing empowerment‑based approaches.

By Jeremy Lucas, Doina Precup
arXiv Machine Learning
Jul 1

End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions

arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.

By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
arXiv Machine Learning
Sep 25

Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes

The paper introduces a vector Bellman theory for multichain robust average‑reward Markov decision processes, addressing the state‑dependent optimal long‑run rewards that arise under uncertainty. It develops a gain‑first, bias‑second optimization principle for finite models with compact, post‑action $(s,a)$‑rectangular ambiguity, yielding a coupled vector gain‑bias system and stationary saddle strategies from all initial states. The authors also characterize solvability conditions, provide certificates for asymptotically affine trajectories of the robust Bellman operator, and design a robust approximately shifted Halpern planning algorithm that converges to the optimal gain vector and produces average‑optimal greedy controllers.

By Yue Wang, George Atia
arXiv AI
Jul 21

A Causal Markov Condition for Value

arXiv:2607. 16717v1 Announce Type: cross Abstract: This paper proposes a causal independence principle for value -- the value Causal Markov Condition (v-CMC) -- and develops the conceptual and mathematical foundations of a "causal value theory" linking causality and utility.

By Olav Benjamin Vassend
arXiv AI
Jun 2

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.

By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv Machine Learning
Jul 28

Universal Decision Learners

arXiv:2605. 30694v2 Announce Type: replace Abstract: Many theories of decision making -- planning, reinforcement learning, causal intervention, online learning, and game-theoretic equilibrium -- turn local information into globally coherent behavior.

By Sridhar Mahadevan