arXiv:2609.24103v1 Announce Type: new
Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. W...
By Larry Preuett, Qiuyi Zhang, Muhammad Aurangzeb Ahmad
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first st...
arXiv:2601. 18930v4 Announce Type: replace-cross Abstract: We are interested in enabling autonomous agents to learn and reason about systems with hidden states, such as locking mechanisms.
By Seiji Shaw, Travis Manderson, Chad Kessens, Nicholas Roy
arXiv:2609.35874v1 Announce Type: new
Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost state...
By Yaacov Pariente, Vadim Indelman
The paper discusses the agent-centric general value function (ACGVF) framework, which allows an agent to decide both which goal to pursue and when to consider a goal finished, beyond merely selecting actions. It notes that ACGVF assumes full observability, while a prior approach used an internal belief state but required externally supplied goals. The note proposes to unify and extend these methods using hierarchical hidden Markov models (HHMMs).
By Kevin Murphy
The paper introduces Interval POMDP Shielding for agents with imperfect perception, aiming to prevent unsafe actions when sensor readings may be misclassified. By estimating perception uncertainty from finite labeled data, the authors construct confidence intervals and model the system as a finite Interval Partially Observable Markov Decision Process. They propose an algorithm that computes a conservative belief set, enabling a runtime shield that guarantees, with high probability, that any action allowed by the shield meets a specified safety lower bound. Experiments on four case studies demonstrate that this shielding approach outperforms state‑of‑the‑art baselines in safety.
By William Scarbro, Ravi Mangal
The paper introduces the Belief-State Engine (BSE), an inference module that supplies a large language model (LLM) with a Bayesian posterior over hidden states in a partially observable Markov decision process (POMDP). By keeping the raw action‑observation log hidden from the LLM, the BSE ensures the agent behaves as a sound Markov policy on the belief MDP, thereby inheriting classical POMDP optimality guarantees. Experiments on the Tiger POMDP and a red‑team attack‑graph task show that BSE‑augmented agents outperform six baselines in task return, belief calibration, and decision consistency.
By Arnab Chattopadhayay, Debdipta Halder
arXiv:2606. 19729v1 Announce Type: cross Abstract: Planning under uncertainty is an essential capability for autonomous robots.
By Marcus Hoerger, Rishikesh Joshi, Rahul Shome, Ian Manchester, Hanna Kurniawati
arXiv:2604.01024v2 Announce Type: replace
Abstract: We study model-based learning of finite-window policies in tabular partially observable Markov decision processes (POMDPs). A common approach to le...
By Philip Jordan, Maryam Kamgarpour
arXiv:2606. 19729v2 Announce Type: replace-cross Abstract: Planning under uncertainty is an essential capability for autonomous robots.
By Marcus Hoerger, Rishikesh Joshi, Rahul Shome, Ian Manchester, Hanna Kurniawati
arXiv:2606. 15654v1 Announce Type: cross Abstract: Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive.
By Wenjing Tang, Xuanjin Jin, Yuan Liu, Renming Huang, Cewu Lu, Panpan Cai
arXiv:2510. 02149v2 Announce Type: replace Abstract: We introduce Action-Triggered Sporadically Traceable Markov Decision Processes (ATST-MDPs), a reinforcement learning framework for partial observability in which full state observations occur stochastically at each step, with probability determined by the chosen action.
By Alexander Ryabchenko, Wenlong Mou