arXiv Computation and Language

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

arXiv AI
Jun 29

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.

By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha
arXiv AI
Aug 14

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.

By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
arXiv Computer Vision
3d ago

EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI

EgoErrorVQA introduces a new egocentric visual question answering task that evaluates visual agents’ ability to detect procedural errors in everyday activities. The paper presents an evaluator agent built on the Agent2Agent protocol and shows that current models struggle with procedural error recognition. It also proposes Ego-ADR, an Adaptive Decoupled Reasoning framework that improves performance on the task, achieving state‑of‑the‑art results.

By Junlong Li, Junxi Li, Jianjun Gao, Chen Cai, Lap-Pui Chau, Yi Wang
arXiv Computer Vision
Aug 21

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

arXiv:2607. 14497v2 Announce Type: replace Abstract: Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world.

By Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu
arXiv Computer Vision
5d ago

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

OmniAssistBench is a new benchmark for evaluating omni-modal large language models (Omni-LLMs) as real‑time video assistants that actively guide users toward goals. The benchmark addresses the challenge of dynamic interaction paths by providing models with predefined priors from source videos, forcing them to follow the same routes as users. The dataset was constructed by reverse‑engineering existing Internet videos into multi‑turn clips, a process that required over 1,000 expert person‑hours. Results show that proprietary Gemini‑3‑Pro scores 66.4/100 while open‑source Qwen3‑Omni‑Instruct scores 51.2, revealing that current models often give incorrect or incomplete answers, struggle with visual prompts, and fail to maintain context or delay responses until target events.

By Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan
Hugging Face Trending Papers
Jun 3

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granularity, undifferentiated difficulty, limited annotation quality, and pervasive answer ambiguity, leaving them unable to diagnose where current models fail.