arXiv:2610.01210v1 Announce Type: new
Abstract: Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, wh...
By Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao
MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.
By Zijie Zhu, Weiren Cai, Yizhou Wang, Zhenjie Yang, Yide Liu, Jiahao Chen, Guanqi He
arXiv:2608.20308v2 Announce Type: replace
Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...
By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.
By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa
arXiv:2608. 20308v1 Announce Type: new Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps.
By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
arXiv:2606. 09243v1 Announce Type: cross Abstract: Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware.
By Yuan Zeng, Yujia Shi, Tiao Tan, Xingting Li, Yaqi Qin, Zongqing Lu, Wenming Yang, Jing-Hao Xue, Qingmin Liao
arXiv:2609.12155v1 Announce Type: new
Abstract: Understanding person-level bi-manual interactions requires not only detecting hands, but also identifying which two hands belong to the same person and...
By Jonghyun Kim, Junho Roh, Yubin Yoon, Hyotae Lee, Jongkuk Park, Taehwan Hwang, Jaechul Kim, Jungho Lee
arXiv:2609.38615v1 Announce Type: cross
Abstract: Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Ex...
By Hongjia Zhai, Xiyu Zhang, Haoran Zhang, Zhichao Ye, Haomin Liu, Guofeng Zhang, Ian Reid, Xingxing Zuo
Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reasoning from the appearance and dynamics of hands and objects themselves.
arXiv:2609.08636v1 Announce Type: cross
Abstract: Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize...
By Qiaohui Chu, Haoyu Zhang, Meng Liu, Haoxiang Shi, Dongmei Jiang, Liqiang Nie
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training.
arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.
By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha