arXiv AI

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

arXiv:2608. 07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness.

arXiv AI
Sep 18

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.

By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic
arXiv AI
4d ago

ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

arXiv:2609.39665v1 Announce Type: new Abstract: Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This...

By Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
arXiv AI
Sep 17

HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models

HINT-Plan is a new method that integrates human intention prediction into robot task planning by using Vision Language Models to infer high‑level human intentions from third‑person images. These intentions are converted into goal states and combined with hierarchical Scene Graphs to formulate joint task‑planning problems in context‑rich environments. In a photorealistic simulation, HINT-Plan achieved a 69.71% success rate, outperforming baselines by up to 35.29% and reducing functional conflicts.

By Yuchen Liu, Luigi Palmieri, Lujun Li, Radu State, Ilche Georgievski, Marco Aiello
arXiv AI
4d ago

X-Planner: Event-Structured Task Planning for Embodied Intelligence

arXiv:2609.25187v2 Announce Type: replace Abstract: Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems...

By Howard Lu, Shalfun Li, Porter Pan, Cris, Lumen, Cyril, Eric Hu, Lily Li, Maeve Zhang, Rain Sun, Robert Wang, KZ Zheng, Viggo Chen, Tim Ding, Regsis Cheng, YJ Xiao, Kian, Hai Lin, Alan Song, Elise Ma, Gody Li, Victor Yao, Yohann Tang, Ingrid Yu, Jason He, James Wang, Ryan Yu, Ping Yang, Chris Pan, Vincent Chen, Roy Gan, Hao Wang, Qian Wang
arXiv AI
Jun 26

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

arXiv:2606. 27251v1 Announce Type: cross Abstract: Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation.

By Junhao Shi, Zezheng Huai, Siyin Wang, Jia Chen, Yubang Wang, Zhaoye Fei, Hechang Chen, Jingjing Gong, Xipeng Qiu, Yu-Gang Jiang
arXiv Computation and Language
Aug 28

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

MineExplorer is a benchmark designed to assess the open‑world exploration abilities of multimodal large language models (MLLMs) in Minecraft. It filters out tasks that rely heavily on Minecraft‑specific knowledge, organizes tasks into ReAct‑style capabilities, and composes atomic tasks into implicit multi‑hop challenges. A multi‑agent synthesis workflow creates reliable task graphs, sandbox scenes, and rule‑based milestone evaluators, and human evaluation confirms its superiority over a single‑agent baseline. Experiments show that while advanced MLLMs can handle many single‑hop tasks, they struggle with longer trajectories that require coordinating hidden prerequisites, and larger models or different thinking modes do not consistently improve performance.

By Tianjie Ju, Yueqing Sun, Zheng Wu, Wei Zhang, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Gongshen Liu, Zhuosheng Zhang