arXiv:2606. 09243v1 Announce Type: cross Abstract: Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware.
By Yuan Zeng, Yujia Shi, Tiao Tan, Xingting Li, Yaqi Qin, Zongqing Lu, Wenming Yang, Jing-Hao Xue, Qingmin Liao
TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces without tactile sensors. It uses 500 hours of pressure-glove data and a 20-hour TwinTouch-20H dataset where generative models render gloved recordings as bare-hand videos, bridging the appearance gap. The system outperforms previous methods on OakInk2, generalizes to unseen natural bare-hand egocentric videos, and improves as glove supervision increases.
TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces from video. It uses 500 hours of pressure‑glove recordings and hand‑object interaction data, and introduces TwinTouch‑20H, a dataset of 20 hours of paired visual data where generative models render gloved recordings as bare‑hand observations while preserving tactile labels. The system outperforms prior methods on OakInk2, generalizes qualitatively to natural bare‑hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales.
By Danyan Zhou, Jinxuan Lu, Jiawei Lin, Tianxing Chen, Chuqiao Lyu, Wenbo Ding
arXiv:2606. 06872v1 Announce Type: cross Abstract: Estimating hand-surface contact pressure from an egocentric view is crucial for AR/VR devices, robotic imitation, and ergonomic analysis.
By Yuan Zeng, Zilue Gao, Yujia Shi, Zongqing Lu, Wenming Yang, QingMin Liao
arXiv:2608.13014v2 Announce Type: replace
Abstract: Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet...
By Andela Ilic, Rachel Schuchert, Yijing Jiang, Christian Holz
EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.
By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa
arXiv:2610.00451v1 Announce Type: cross
Abstract: Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual...
By Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI)
arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.
By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
arXiv:2608.28386v1 Announce Type: new
Abstract: Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic o...
By Semin Kim, Haechan Shin, Jongyoo Kim
arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.
By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin
arXiv:2606. 24450v1 Announce Type: cross Abstract: Perceiving physical contact is fundamental to dexterous manipulation.
By Soham Patil, Avirup Das, Sourabh Bhosale, Spandan Roy
arXiv:2610.01210v1 Announce Type: new
Abstract: Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, wh...
By Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao