EgoErrorVQA introduces a new egocentric visual question answering task that evaluates visual agents’ ability to detect procedural errors in everyday activities. The paper presents an evaluator agent built on the Agent2Agent protocol and shows that current models struggle with procedural error recognition. It also proposes Ego-ADR, an Adaptive Decoupled Reasoning framework that improves performance on the task, achieving state‑of‑the‑art results.
By Junlong Li, Junxi Li, Jianjun Gao, Chen Cai, Lap-Pui Chau, Yi Wang
arXiv:2609.14473v1 Announce Type: new
Abstract: Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical...
By Avijit Dasgupta, Shayon Dasgupta, Zakaria Laskar, C. V. Jawahar, Karteek Alahari
arXiv:2607. 04162v1 Announce Type: cross Abstract: Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures.
By Iok Tong Lei, QianZhi Li, Ying Jie Yap, Yujie Zhang, Rui Zhong, Haichao Gui, Xiaolong Liu, Zhidong Deng
The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.
By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
arXiv:2607. 26865v1 Announce Type: cross Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems.
By Amirmohammad Farzaneh, Osvaldo Simeone
The majority of our everyday activities are procedural and consist of sequences of interdependent steps. However, existing benchmarks for Visual Agents and Visual Language Models (VLMs) overlook the e...