arXiv:2503. 22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition.
By Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding
arXiv:2608. 13415v1 Announce Type: cross Abstract: We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks.
By Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, George Konidaris
arXiv:2606. 03312v1 Announce Type: cross Abstract: While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robots are expected to choose actions that prioritize other values than task success, such as human autonomy, efficiency, or social appropriateness.
By Jongwook Han, Hyeongjin Kim, Yohan Jo
Physical Agentic AI proposes an architecture that links semantic planning with physical execution for robot crews. Each robot exposes a typed skill library, while a foundation model planner decomposes tasks into phases and assigns robot‑skill pairs. A Robot Orchestrator validates and authorizes one skill at a time, ensuring actions are grounded in robot capabilities, system state, and workflow constraints before actuation.
By Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake
The paper introduces a method for generating legible plans in arbitrary PDDL domains by extending prior legibility research to classical planning without custom planners. It incorporates a second‑order theory of mind to estimate the observer’s perspective, enabling robots to implicitly communicate goals in human‑robot teaming. Benchmark results show that increasing legibility typically trades off with plan efficiency, and a regularizing factor is needed to balance the two.
By Michele Persiani, Thomas Hellstr\"om
The paper introduces STEP, a State‑Aware Task Estimator and Planner that uses multi‑modal large language models to explicitly estimate system states and predict state transitions during task planning. By forecasting future states alongside actions, STEP reduces hallucinated actions and improves task‑convergent planning. In a simulated robot assembly task, STEP outperforms the state‑of‑the‑art by 32.8% in action executability and 14.8% in final‑state error.
By Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba
arXiv:2603.08814v2 Announce Type: replace-cross
Abstract: Long-horizon task planning for heterogeneous multi-robot systems is essential for deploying collaborative teams in real-world environments; y...
By Piyush Gupta, Sangjae Bae, Jiachen Li, David Isele
arXiv:2608. 05588v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones.
By He Jiang, Jingtian Yan, Yulun Zhang, Yimin Tang, Tanishq Duhan, Rishi Veerapaneni, Guillaume Sartoretti, Jiaoyang Li
The paper introduces Disjoint Parameter Training (DPT), a framework that addresses Skill Conflict—where shared encoder parameters hinder separate tasks of motion prediction and safety planning—by training tasks on distinct parameter subsets before merging. DPT employs sparse merging to integrate only the most influential parameters, reducing interference and enhancing representational capacity. Experiments on JRDB and JTA benchmarks show that DPT outperforms existing unified models, demonstrating its effectiveness for safe, resource‑efficient robot navigation.
By Taewon Seo, Seonae Jeon, Giwon Lee, Kuk-Jin Yoon, Daehee Park
arXiv:2606. 16480v1 Announce Type: cross Abstract: Robots deployed in the real world must plan motions across diverse scenarios without per-scenario retuning.
By Youngjae Min, Jovin D'sa, Faizan M. Tariq, David Isele, Navid Azizan, Sangjae Bae
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution.
arXiv:2609.25351v1 Announce Type: cross
Abstract: We focus on human-robot collaborative transport, a challenging task of broad relevance spanning logistics, manufacturing, and the home, in which a us...
By Elvin Yang, Christoforos Mavrogiannis