arXiv Computer Vision

G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

arXiv:2608. 20157v1 Announce Type: new Abstract: Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets.

arXiv AI
Jun 29

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.

By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha
arXiv Computer Vision
Aug 27

Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

The paper introduces a redundancy-aware fusion framework for EgoExo proficiency estimation, which integrates fine-grained motion cues from egocentric views with spatial context from exocentric views. It identifies multiview redundancy and overfitting as key challenges and proposes two modules—AdaMVS for adaptive view selection and VIB-GB for compressing redundant signals—to address them. Experiments on EgoExo-4D and EgoExo-Fitness show that the method learns to select informative views and fuse them effectively, achieving state‑of‑the‑art results.

By Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert
Hugging Face Trending Papers
Jul 9

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reasoning from the appearance and dynamics of hands and objects themselves.

arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
arXiv AI
Jul 2

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

arXiv:2607. 00547v1 Announce Type: cross Abstract: Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation.

By Jihyeok Jung (KAIST AI), Jeewu Lee (Sogang University), Sanghyeop Kim (Sogang University), Chanhee Han (Ministry of Science and ICT), Seong Joon Oh (KAIST AI)