EgoErrorVQA introduces a new egocentric visual question answering task that evaluates visual agents’ ability to detect procedural errors in everyday activities. The paper presents an evaluator agent built on the Agent2Agent protocol and shows that current models struggle with procedural error recognition. It also proposes Ego-ADR, an Adaptive Decoupled Reasoning framework that improves performance on the task, achieving state‑of‑the‑art results.
By Junlong Li, Junxi Li, Jianjun Gao, Chen Cai, Lap-Pui Chau, Yi Wang
arXiv:2607. 24770v1 Announce Type: new Abstract: Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing physical actions.
By Azizul Zahid, Subrata Biswas, Bashima Islam, Sai Swaminathan
EgoArgus is a new, human‑annotated dataset that tests visual‑language models (VLMs) as situational assistants in five everyday dialogue‑video scenarios. It evaluates how well VLMs understand and decide when to intervene, especially when visual and textual cues are helpful, irrelevant, or conflicting. The study finds that current VLMs still struggle to reliably act as egocentric assistants and that existing modality‑bias mitigation methods offer limited improvement.
By Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen
arXiv:2609.22942v1 Announce Type: new
Abstract: Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plau...
By Zhenchen Tang, Bo Peng, Zichuan Wang, Songlin Yang, Leilei Cao, Fengjie Zhu, Jing Dong
arXiv:2609.14473v1 Announce Type: new
Abstract: Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical...
By Avijit Dasgupta, Shayon Dasgupta, Zakaria Laskar, C. V. Jawahar, Karteek Alahari
arXiv:2606. 13929v1 Announce Type: cross Abstract: Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored.
By Yijun Liang, Hengguang Zhou, Ming Li, Lichen Li, Cho-Jui Hsieh, Tianyi Zhou
Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.
By Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
While Large Language Models (LLMs) excel as static solvers, transforming them into autonomous agents remains challenging. This transition requires continuous environmental interaction, yet current agents lack the necessary persistent procedural memory.
BuddyVQA is a new benchmark for companion‑style question answering on egocentric streaming video, comprising 21.6K questions tied to 6K highlight moments across 1,012 long first‑person videos. It emphasizes two often overlooked aspects of daily first‑person QA: ego‑deictic expressions and interactively chained questions, requiring models to resolve visual pronouns and infer user intent within a long‑form streaming context. The authors propose MyBuddy, a multimodal chain‑of‑thought QA assistant that uses a question filter and multi‑level memory to efficiently retrieve visual and QA information, achieving significant performance gains on BuddyVQA and generalizing to other streaming and common video QA benchmarks.
By Hangyu Qin, Junbin Xiao, Shenglang Zhang, Angela Yao
arXiv:2505. 23399v2 Announce Type: replace Abstract: We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning.
By Jusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen, Haoyi Jiang, Wenhao Chai, Jian Wang, Keze Wang
arXiv:2606. 29824v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel as static solvers, transforming them into autonomous agents remains challenging.
By Chengfeng Zhao, Yuqiao Tan, Shizhu He, Yequan Wang, Jun Zhao, Kang Liu
arXiv:2606. 05275v1 Announce Type: cross Abstract: We study the personal camera roll visual question answering setting.
By Thao Nguyen, Krishna Kumar Singh, Donghyun Kim, Yong Jae Lee, Yuheng Li