The paper introduces a new paradigm called VLM-as-probabilistic-grounder, which models the uncertainty of Vision‑Language Model (VLM) predicate groundings as a probability distribution over symbolic states. This probabilistic grounding allows belief‑space planning, producing more robust plans in partially observable settings. Experiments in simulated household robot environments demonstrate that this approach improves robustness and task success compared to deterministic grounding methods.
By Guy Azran, Michael Navat, Sarah Keren
arXiv:2505.13180v3 Announce Type: replace
Abstract: Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works ex...
By Matteo Merler, Nicola Dainese, Minttu Alakuijala, Giovanni Bonetta, Pietro Ferrazzi, Yu Tian, Bernardo Magnini, Pekka Marttinen
arXiv:2606. 15654v1 Announce Type: cross Abstract: Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive.
By Wenjing Tang, Xuanjin Jin, Yuan Liu, Renming Huang, Cewu Lu, Panpan Cai
arXiv:2607. 08024v1 Announce Type: cross Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility.
By Emily Jin, Joy Hsu, Yiqing Xu, Weiyu Liu, Nick Haber, Jiajun Wu
HINT-Plan is a new method that integrates human intention prediction into robot task planning by using Vision Language Models to infer high‑level human intentions from third‑person images. These intentions are converted into goal states and combined with hierarchical Scene Graphs to formulate joint task‑planning problems in context‑rich environments. In a photorealistic simulation, HINT-Plan achieved a 69.71% success rate, outperforming baselines by up to 35.29% and reducing functional conflicts.
By Yuchen Liu, Luigi Palmieri, Lujun Li, Radu State, Ilche Georgievski, Marco Aiello
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning.
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executab...
arXiv:2608. 20084v1 Announce Type: cross Abstract: Robots executing long-horizon manipulation tasks from natural-language instructions must reason about both semantic task structure and geometric feasibility.
By Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar, Edgar Simo-Serra
arXiv:2609.37554v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructi...
By {\L}ukasz Sobczak, Nur Kele\c{s}o\u{g}lu, S{\l}awomir Piotr Nowak
arXiv:2608. 16794v1 Announce Type: cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities.
By Mohammad Albinhassan, Yuming Feng, Alessandra Russo, Pranava Madhyastha
arXiv:2606. 12910v1 Announce Type: cross Abstract: For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time.
By Allison Andreyev, Landon Eum, Nestor Tiglao, Romel Gomez
The paper introduces OHCAM, an online method for learning action models that include conditional and quantified effects from limited interactions. It maintains a belief over possible models and actively chooses actions that maximize disagreement among hypotheses to reduce uncertainty, while handling noisy observations. Starting with simple hypotheses, OHCAM expands complexity only when necessary, achieving sample‑efficient learning that outperforms baselines on benchmark domains and is validated on a Kinova Gen3 robot.
By Jeffrey Jewett, William Solow, Sandhya Saisubramanian