arXiv:2609.08636v1 Announce Type: cross
Abstract: Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize...
By Qiaohui Chu, Haoyu Zhang, Meng Liu, Haoxiang Shi, Dongmei Jiang, Liqiang Nie
The paper introduces UniWM, a unified, memory‑augmented world model that merges egocentric visual foresight and planning into a single multimodal autoregressive backbone. By grounding action selection in visually imagined outcomes and using a hierarchical memory to fuse short‑term perception with long‑term trajectory context, UniWM aligns prediction with control and improves navigation stability. Experiments on four challenging benchmarks and the 1X Humanoid Dataset show up to 30% higher success rates, reduced trajectory errors, zero‑shot generalization to unseen datasets, and scalability to high‑dimensional humanoid navigation.
By Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng
arXiv:2606.08495v2 Announce Type: replace-cross
Abstract: Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specif...
By Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang, Kun Li, Kai Chen
arXiv:2609.06195v1 Announce Type: new
Abstract: This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds...
By Azusa Sawada, Allan Wang, Hideo Saito, Aaron Steinfeld
arXiv:2608.29927v1 Announce Type: new
Abstract: We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each indivi...
By Daeyun Shin, Yunhan Zhao, Shu Kong, Alexander C. Berg, Charless Fowlkes
The paper introduces HAP, a Hand-Driven Active Perception framework that predicts future six‑degree‑of‑freedom head motion in egocentric settings by conditioning on observed hand motion and inferred target context. HAP constructs a Predictive Target‑Centric Amodal Occlusion Graph to model current and potential occlusions among candidate objects, fuses this with hand and head motion history, and blends the learned trajectory with a constant‑velocity prior. Experiments on a public dataset and a newly released Bottle RGB‑D dataset demonstrate that HAP outperforms baseline methods in head‑motion prediction, highlighting the importance of hand‑driven intention and dynamic occlusion reasoning.
By Yunji Feng, Junyi Ma, Guanzhong Sun, Chenyang Xu, Hesheng Wang