arXiv Machine Learning By Zhiwen Qiu, Ziang Liu, Wenqian Niu, Tapomayukh Bhattacharjee, Saleh Kalantari

EgoCogNav: Cognition-aware Human Egocentric Navigation

Read the original on arXiv Machine Learning →

arXiv:2511. 17581v3 Announce Type: replace Abstract: Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight

The paper introduces UniWM, a unified, memory‑augmented world model that merges egocentric visual foresight and planning into a single multimodal autoregressive backbone. By grounding action selection in visually imagined outcomes and using a hierarchical memory to fuse short‑term perception with long‑term trajectory context, UniWM aligns prediction with control and improves navigation stability. Experiments on four challenging benchmarks and the 1X Humanoid Dataset show up to 30% higher success rates, reduced trajectory errors, zero‑shot generalization to unseen datasets, and scalability to high‑dimensional humanoid navigation.

By Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng
arXiv Computer Vision
Sep 1

Everybody Tracking Every Body

arXiv:2608.29927v1 Announce Type: new Abstract: We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each indivi...

By Daeyun Shin, Yunhan Zhao, Shu Kong, Alexander C. Berg, Charless Fowlkes
arXiv Computer Vision
Sep 17

HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction

The paper introduces HAP, a Hand-Driven Active Perception framework that predicts future six‑degree‑of‑freedom head motion in egocentric settings by conditioning on observed hand motion and inferred target context. HAP constructs a Predictive Target‑Centric Amodal Occlusion Graph to model current and potential occlusions among candidate objects, fuses this with hand and head motion history, and blends the learned trajectory with a constant‑velocity prior. Experiments on a public dataset and a newly released Bottle RGB‑D dataset demonstrate that HAP outperforms baseline methods in head‑motion prediction, highlighting the importance of hand‑driven intention and dynamic occlusion reasoning.

By Yunji Feng, Junyi Ma, Guanzhong Sun, Chenyang Xu, Hesheng Wang