Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes
Read the original on arXiv Machine Learning →The paper introduces a vector Bellman theory for multichain robust average‑reward Markov decision processes, addressing the state‑dependent optimal long‑run rewards that arise under uncertainty. It develops a gain‑first, bias‑second optimization principle for finite models with compact, post‑action $(s,a)$‑rectangular ambiguity, yielding a coupled vector gain‑bias system and stationary saddle strategies from all initial states. The authors also characterize solvability conditions, provide certificates for asymptotically affine trajectories of the robust Bellman operator, and design a robust approximately shifted Halpern planning algorithm that converges to the optimal gain vector and produces average‑optimal greedy controllers.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.