The paper "Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences" addresses the challenge of natural-language instructions that omit details needed for embodied action. It introduces the Preference-based Planning (PbP) benchmark, comprising 5,000 evaluation groups and 290 preferences across three levels, to systematically evaluate agents’ ability to infer latent user preferences from a few demonstrations. The authors propose the two-stage Inferring the Unspoken (InTU) framework, which first verbalizes inferred preferences from multimodal demonstrations and then generates action plans conditioned on that explicit representation, showing that explicit verbalization improves alignment and robustness compared to direct end-to-end planning.
By Manjie Xu, Xinyi Yang, Wei Liang, Chi Zhang, Yixin Zhu
The paper introduces STEP, a State‑Aware Task Estimator and Planner that uses multi‑modal large language models to explicitly estimate system states and predict state transitions during task planning. By forecasting future states alongside actions, STEP reduces hallucinated actions and improves task‑convergent planning. In a simulated robot assembly task, STEP outperforms the state‑of‑the‑art by 32.8% in action executability and 14.8% in final‑state error.
By Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba
arXiv:2606. 27826v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requires not only achieving instructed goals but also acting in socially appropriate ways.
By Shiyun Zhao, Xinwei Song, Tianyu Guo, Xiaomeng Gao, Mingyuan Liu, Xu Han, Yuanyuan Zhang, Zhenliang Zhang, Xue Feng, Bo Dai
arXiv:2607. 13621v1 Announce Type: new Abstract: Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode.
By Kun Yu, Jianhua Yang, Yixiang Chen, Changwei Wang, Hongyuan Yu, Yan Huang, Fushuo Huo, Ya Jing, Zhumin Chen, Keji He
arXiv:2606.01063v3 Announce Type: replace
Abstract: Theory-of-Mind (ToM) reasoning enables embodied agents to understand human beliefs, goals, and intentions, but existing benchmarks mainly evaluate...
By Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Wen-Huang Cheng, Jianlong Fu
arXiv:2606. 01810v1 Announce Type: new Abstract: Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning.
By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Heng Cao, Zirui Song, Yifan Yang, Chong Luo, Bei Liu, Yiming Li