arXiv Computer Vision

DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction

arXiv Computer Vision
Aug 28

Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions

Grasp in Gaussians (GraG) is a fast, robust method for reconstructing dynamic 3D hand‑object interactions from a single monocular video. It leverages pretrained hand and object priors and represents the scene with a compact Sum‑of‑Gaussians (SoG) model, enabling efficient tracking while preserving geometric fidelity. Experiments show GraG achieves temporally coherent reconstructions on long sequences 4.4×–38.9× faster than prior work.

By Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, Christian Theobalt
arXiv AI
Jun 9

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.

By Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan, Shiyu Gao, Ying Ai, Hui Xu, Yongqi Hu, Hengyi Zhang, Yang Xie, Zhaxizhuoma, Yue Zhao, Zhenshan Bing, Yan Ding, Jianxing Liu
arXiv AI
2d ago

PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

arXiv:2610.00451v1 Announce Type: cross Abstract: Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual...

By Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI)
Hugging Face Trending Papers
Sep 17

TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces without tactile sensors. It uses 500 hours of pressure-glove data and a 20-hour TwinTouch-20H dataset where generative models render gloved recordings as bare-hand videos, bridging the appearance gap. The system outperforms previous methods on OakInk2, generalizes to unseen natural bare-hand egocentric videos, and improves as glove supervision increases.