arXiv:2606. 27826v3 Announce Type: replace Abstract: Embodied agents driven by multimodal large language models (MLLMs) can often complete everyday tasks from visual observations, but goal achievement does not establish whether they proactively respect unstated social norms.
By Shiyun Zhao, Xinwei Song, Tianyu Guo, Xiaomeng Gao, Mingyuan Liu, Xu Han, Yuanyuan Zhang, Zhenliang Zhang, Xue Feng, Bo Dai
The paper "Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences" addresses the challenge of natural-language instructions that omit details needed for embodied action. It introduces the Preference-based Planning (PbP) benchmark, comprising 5,000 evaluation groups and 290 preferences across three levels, to systematically evaluate agents’ ability to infer latent user preferences from a few demonstrations. The authors propose the two-stage Inferring the Unspoken (InTU) framework, which first verbalizes inferred preferences from multimodal demonstrations and then generates action plans conditioned on that explicit representation, showing that explicit verbalization improves alignment and robustness compared to direct end-to-end planning.
By Manjie Xu, Xinyi Yang, Wei Liang, Chi Zhang, Yixin Zhu
The paper introduces FISER, a framework that explicitly infers human goals and intentions before planning actions for AI agents to follow natural language instructions in collaborative embodied tasks. It employs Transformer-based models and is evaluated on the HandMeThat benchmark, outperforming end-to-end approaches and strong baselines such as Chain of Thought prompting. FISER achieves state‑of‑the‑art performance on this embodied social reasoning task.
By Yanming Wan, Yue Wu, Yiping Wang, Jiayuan Mao, Natasha Jaques
arXiv:2609.08602v1 Announce Type: new
Abstract: Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequenc...
By Tianyi Ma, Parisa Kordjamshidi
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executab...
arXiv:2603.26741v2 Announce Type: replace-cross
Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, languag...
By Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Yetong Sha, Qiyu Hu, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi Cheng
The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.
By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
The paper introduces Instruct-to-Act, a system that decouples planning and control by combining a vision‑language model (VLM) planner with a world‑model controller. The VLM generates sparse, high‑level text instructions, while the controller executes them at high frequency, trained via relabeling rollouts with synthetic instructions and joint optimization of behavior cloning, reward, and world‑model objectives. Across seven embodied environments—including multi‑agent settings—this approach outperforms controller‑only and direct VLM action methods, maintains fast control, and allows swapping pretrained VLM planners without fine‑tuning, achieving competitive results with strong baselines on most tasks.
By Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr
arXiv:2608. 06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents.
By Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen
arXiv:2609.39235v1 Announce Type: cross
Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
By Ali Alrasheed, Basim Azam, Naveed Akhtar
The paper introduces ProAction, a multimodal dataset of 10,000 samples comprising visual, audio, and text inputs across 12 daily-life scenarios, designed to support the Proactive Robot Action Reasoning (ProRobo) problem. It presents a two-stage human-in-the-loop annotation pipeline that incorporates appraisal and Theory-of-Mind considerations to generate cognitively grounded high-level action labels. The authors benchmark multimodal large language models and propose MMC2Act, showing that training on ProAction significantly improves proactive action reasoning compared to general-purpose models.
By Zhihao Gu, Kechao Zhu, Yuanfeng Wu, Mohan Liu, Ankit Kumar Shaw, ChenDong Hong, Xuanyu Chen, Dengchen Mei, Xu Tianyi, Lin Wang
The paper introduces Instruct-to-Act, a system that decouples high‑level planning from low‑latency control by combining a vision‑language model (VLM) planner with a world‑model controller. The VLM generates sparse, high‑level text instructions, while the controller executes them autonomously at high frequency. Experiments across seven embodied environments, including multi‑agent settings, show that this approach outperforms both controller‑only and direct VLM action‑generation methods, maintains fast control, and allows swapping in different pretrained VLM planners without fine‑tuning.