arXiv Machine Learning

EgoCogNav: Cognition-aware Human Egocentric Navigation

arXiv:2511. 17581v3 Announce Type: replace Abstract: Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding.

arXiv AI
Aug 19

Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight

The paper introduces UniWM, a unified, memory‑augmented world model that merges egocentric visual foresight and planning into a single multimodal autoregressive backbone. By grounding action selection in visually imagined outcomes and using a hierarchical memory to fuse short‑term perception with long‑term trajectory context, UniWM aligns prediction with control and improves navigation stability. Experiments on four challenging benchmarks and the 1X Humanoid Dataset show up to 30% higher success rates, reduced trajectory errors, zero‑shot generalization to unseen datasets, and scalability to high‑dimensional humanoid navigation.

By Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng
arXiv Computer Vision
Sep 1

Everybody Tracking Every Body

arXiv:2608.29927v1 Announce Type: new Abstract: We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each indivi...

By Daeyun Shin, Yunhan Zhao, Shu Kong, Alexander C. Berg, Charless Fowlkes
arXiv Computer Vision
Sep 17

HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction

The paper introduces HAP, a Hand-Driven Active Perception framework that predicts future six‑degree‑of‑freedom head motion in egocentric settings by conditioning on observed hand motion and inferred target context. HAP constructs a Predictive Target‑Centric Amodal Occlusion Graph to model current and potential occlusions among candidate objects, fuses this with hand and head motion history, and blends the learned trajectory with a constant‑velocity prior. Experiments on a public dataset and a newly released Bottle RGB‑D dataset demonstrate that HAP outperforms baseline methods in head‑motion prediction, highlighting the importance of hand‑driven intention and dynamic occlusion reasoning.

By Yunji Feng, Junyi Ma, Guanzhong Sun, Chenyang Xu, Hesheng Wang
arXiv AI
2d ago

NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video

NextMe-800 is an approximately 800‑hour first‑person video dataset collected from a single volunteer over 126 days, featuring 1 Hz images, gaze, and audio. The data are captioned at five hierarchical abstraction levels—from atomic actions to major activities—enabling personalized action anticipation as an open‑vocabulary K‑step sequence prediction task. The authors also introduce NextAct, a 1,500‑point benchmark that combines NextMe‑800 with the multi‑person EgoLife dataset, and evaluate models using an embedding‑based soft edit distance to assess how well personal behavior can be anticipated across abstraction levels and prediction horizons.

By Zhaoxu Meng, Yiming Sun, Mingyuan Gao, Jiachang Zhang, Zhuhan Dai, Yipeng Du, Zheng Lian, Jian-Qiao Zhu
arXiv Computer Vision
Aug 27

Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

The paper introduces a redundancy-aware fusion framework for EgoExo proficiency estimation, which integrates fine-grained motion cues from egocentric views with spatial context from exocentric views. It identifies multiview redundancy and overfitting as key challenges and proposes two modules—AdaMVS for adaptive view selection and VIB-GB for compressing redundant signals—to address them. Experiments on EgoExo-4D and EgoExo-Fitness show that the method learns to select informative views and fuse them effectively, achieving state‑of‑the‑art results.

By Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert
arXiv Computer Vision
Sep 7

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.

By Zijie Zhu, Weiren Cai, Yizhou Wang, Zhenjie Yang, Yide Liu, Jiahao Chen, Guanqi He