arXiv AI By Yujiao Chen

Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States

Read the original on arXiv AI →

arXiv:2606. 00970v1 Announce Type: new Abstract: We study risk-neutral control in Markov decision processes with an absorbing catastrophic state.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin