arXiv:2606. 27757v1 Announce Type: new Abstract: Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns regarding robustness and reliability.
By Jiajing Zhang, Jiamei Jiang, Chenyang Zhang, Feifei Mo, Linjing Li, Daniel Zeng
arXiv:2608. 09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose.
By Rohan Bhagra, Mahantesh Halapannavar, Uddhav Bhattarai
The paper introduces a new paradigm called VLM-as-probabilistic-grounder, which models the uncertainty of Vision‑Language Model (VLM) predicate groundings as a probability distribution over symbolic states. This probabilistic grounding allows belief‑space planning, producing more robust plans in partially observable settings. Experiments in simulated household robot environments demonstrate that this approach improves robustness and task success compared to deterministic grounding methods.
By Guy Azran, Michael Navat, Sarah Keren
The paper introduces STEP, a State‑Aware Task Estimator and Planner that uses multi‑modal large language models to explicitly estimate system states and predict state transitions during task planning. By forecasting future states alongside actions, STEP reduces hallucinated actions and improves task‑convergent planning. In a simulated robot assembly task, STEP outperforms the state‑of‑the‑art by 32.8% in action executability and 14.8% in final‑state error.
By Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba
GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.
By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic
The paper introduces CLUE, a framework that lets robots actively resolve contextual uncertainty for underspecified natural language tasks. CLUE employs an LLM-derived policy to generate task-relevant hypotheses and plans, then uses an online language-embedded map to ground these into actions, refining its plan through closed-loop interaction. Experiments on a Boston Dynamics Spot across diverse indoor and outdoor settings show CLUE achieving near-oracle performance and outperforming LLM planners without closed-loop feedback by a significant margin.
By Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas, Vijay Kumar
arXiv:2510. 00182v2 Announce Type: replace-cross Abstract: While we know that large language models (LLMs) can solve some planning problems, we do not understand the extent of these capabilities for robotics.
By Jorge Mendez-Mendez
In partially observable settings, agents must act without full knowledge of the world state and rely on uncertain state-estimation pipelines. Obtaining grounded and verifiable symbolic plans under suc...
arXiv:2606. 04226v1 Announce Type: cross Abstract: Simulation environments are useful for both robot policy learning and planning verification and validation.
By Charlie Gauthier, Sacha Morin, Liam Paull
The paper introduces an LLM chaining architecture for General Purpose Service Robots that splits instruction classification and action generation into two stages, cutting prompt length by about 45% and boosting planning consistency. Evaluation on 100 synthetic GPSR commands across three language models shows consistent improvements over single-prompt methods, with up to +37 percentage points gain on local models. Real‑robot trials on the Toyota HSR confirm that while planning success improves, execution-layer failures remain the main obstacle to full task completion.
By Lucas Da Mota Bruno, Jiahao Sim, Yoshinobu Hagiwara
CT‑SAFR is a multi‑layered verification framework designed to enhance the safety and faithfulness of Chain‑of‑Thought reasoning in autonomous robots. The framework achieves a 94.2% hallucination detection rate with sub‑500 ms latency, and a warehouse robot case study shows an 87% reduction in unsafe reasoning outputs. The study also offers recommendations for responsible deployment of reasoning‑capable autonomous robots.
By Cagri Temel
The paper introduces a ROS-Agent architecture that enhances task reliability and execution efficiency for open‑source LLM‑powered robotic agents. It adds a MetaTool that forces the LLM to produce a structured pseudo‑code plan before any action, storing this plan in a scratchpad to separate planning from execution. Experiments on a custom mobile robot show up to ~24% improvement in complex task completion and contextual consistency compared to the baseline.
By Kazi Abrar Mahmud, Nilotpaul Kundu Dhurubo, Tamal Kirttonia, Sabbir Hossain Ujjal, Mohammad Ariful Haque