arXiv Computer Vision

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

arXiv Computer Vision
Sep 7

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.

By Zijie Zhu, Weiren Cai, Yizhou Wang, Zhenjie Yang, Yide Liu, Jiahao Chen, Guanqi He
arXiv Machine Learning
Jun 10

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations

arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.

By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin
arXiv AI
Oct 2

PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

PACT is an end‑to‑end model that jointly learns human pose, contacts, and contact forces from monocular video. It augments a human reconstruction foundation model with learnable contact‑force tokens and a temporal transformer, and uses physics‑based supervision to enforce consistency between motion and forces. The authors also create a data annotation pipeline and a real‑world climbing benchmark, ForceWall, to train and evaluate the system, achieving state‑of‑the‑art performance and better generalization than staged approaches.

By Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI)
arXiv AI
Jun 9

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.

By Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan, Shiyu Gao, Ying Ai, Hui Xu, Yongqi Hu, Hengyi Zhang, Yang Xie, Zhaxizhuoma, Yue Zhao, Zhenshan Bing, Yan Ding, Jianxing Liu
arXiv Computer Vision
Sep 16

EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset

EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.

By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa
Hugging Face Trending Papers
Aug 12

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training.

arXiv Computer Vision
Sep 1

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

arXiv:2608.20308v2 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...

By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
arXiv AI
Sep 18

TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces from video. It uses 500 hours of pressure‑glove recordings and hand‑object interaction data, and introduces TwinTouch‑20H, a dataset of 20 hours of paired visual data where generative models render gloved recordings as bare‑hand observations while preserving tactile labels. The system outperforms prior methods on OakInk2, generalizes qualitatively to natural bare‑hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales.

By Danyan Zhou, Jinxuan Lu, Jiawei Lin, Tianxing Chen, Chuqiao Lyu, Wenbo Ding