arXiv AI
Sep 16

Bridging Learned Visual Perception and Symbolic Belief-Space Planning

The paper introduces a new paradigm called VLM-as-probabilistic-grounder, which models the uncertainty of Vision‑Language Model (VLM) predicate groundings as a probability distribution over symbolic states. This probabilistic grounding allows belief‑space planning, producing more robust plans in partially observable settings. Experiments in simulated household robot environments demonstrate that this approach improves robustness and task success compared to deterministic grounding methods.

By Guy Azran, Michael Navat, Sarah Keren
arXiv AI
Jun 16

PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty

arXiv:2606. 15654v1 Announce Type: cross Abstract: Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive.

By Wenjing Tang, Xuanjin Jin, Yuan Liu, Renming Huang, Cewu Lu, Panpan Cai
arXiv AI
Sep 17

HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models

HINT-Plan is a new method that integrates human intention prediction into robot task planning by using Vision Language Models to infer high‑level human intentions from third‑person images. These intentions are converted into goal states and combined with hierarchical Scene Graphs to formulate joint task‑planning problems in context‑rich environments. In a photorealistic simulation, HINT-Plan achieved a 69.71% success rate, outperforming baselines by up to 35.29% and reducing functional conflicts.

By Yuchen Liu, Luigi Palmieri, Lujun Li, Radu State, Ilche Georgievski, Marco Aiello
Hugging Face Trending Papers
Aug 17

Neurosymbolic Embodied Agents

Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning.