Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
arXiv:2607. 11523v1 Announce Type: cross Abstract: When should an intelligent assistant speak up without being asked?
arXiv:2607. 11523v1 Announce Type: cross Abstract: When should an intelligent assistant speak up without being asked?
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive.
arXiv:2604. 08342v2 Announce Type: replace Abstract: Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains.
arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.
arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.
arXiv:2606. 17627v1 Announce Type: cross Abstract: Fine-grained action recognition in egocentric video is challenging for Vision-Language Models (VLMs): actions often differ only in small visual cues, and a single model tends to be biased toward a subset of these cues.
arXiv:2606. 04806v1 Announce Type: cross Abstract: LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior.
EgoErrorVQA introduces a new egocentric visual question answering task that evaluates visual agents’ ability to detect procedural errors in everyday activities. The paper presents an evaluator agent built on the Agent2Agent protocol and shows that current models struggle with procedural error recognition. It also proposes Ego-ADR, an Adaptive Decoupled Reasoning framework that improves performance on the task, achieving state‑of‑the‑art results.
arXiv:2607. 14497v2 Announce Type: replace Abstract: Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world.
OmniAssistBench is a new benchmark for evaluating omni-modal large language models (Omni-LLMs) as real‑time video assistants that actively guide users toward goals. The benchmark addresses the challenge of dynamic interaction paths by providing models with predefined priors from source videos, forcing them to follow the same routes as users. The dataset was constructed by reverse‑engineering existing Internet videos into multi‑turn clips, a process that required over 1,000 expert person‑hours. Results show that proprietary Gemini‑3‑Pro scores 66.4/100 while open‑source Qwen3‑Omni‑Instruct scores 51.2, revealing that current models often give incorrect or incomplete answers, struggle with visual prompts, and fail to maintain context or delay responses until target events.
Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granularity, undifferentiated difficulty, limited annotation quality, and pervasive answer ambiguity, leaving them unable to diagnose where current models fail.
arXiv:2606. 00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos.