arXiv:2605.03927v3 Announce Type: replace
Abstract: Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle...
By Xiaowen Sun, Matthias Kerzel, Mengdi Li, Xufeng Zhao, Paul Striker, Stefan Wermter
arXiv:2508. 08983v2 Announce Type: replace-cross Abstract: Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new constraints.
By Ben Zandonati, Tom\'as Lozano-P\'erez, Leslie Pack Kaelbling
arXiv:2609.39665v1 Announce Type: new
Abstract: Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This...
By Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.
By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
arXiv:2607.05780v2 Announce Type: replace-cross
Abstract: While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same functio...
By Chuhao Zhou, Liquan Wang, Shuxin Cao, Xiangyu Chen, Yuxuan Hu, Boyu Ma, Animesh Garg, Jianfei Yang
arXiv:2609.22332v1 Announce Type: cross
Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act...
By Jiadi You, Qize Yu, Yue Chen, Minghong Cai, Zhide Zhong, Yuran Wang, Bowen Ping, Jiaqi Liang, Zhenhao Shen, Haodong Yan, Yinchuan Li, Ruihai Wu, Xiaojuan Qi, Yingcong Chen