arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.
By Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan, Shiyu Gao, Ying Ai, Hui Xu, Yongqi Hu, Hengyi Zhang, Yang Xie, Zhaxizhuoma, Yue Zhao, Zhenshan Bing, Yan Ding, Jianxing Liu
arXiv:2604. 01001v2 Announce Type: replace-cross Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation.
By Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu, Jiafei Cao, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, Xudong Xu
arXiv:2606. 16202v1 Announce Type: cross Abstract: Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformable dynamics, such as elastic materials and fabrics, remains a major challenge for computer vision and robotics.
By Hyunjin Kim, Ri-Zhao Qiu, Guangqi Jiang, Xiaolong Wang
MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.
By Zijie Zhu, Weiren Cai, Yizhou Wang, Zhenjie Yang, Yide Liu, Jiahao Chen, Guanqi He
arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.
By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
arXiv:2606. 09243v1 Announce Type: cross Abstract: Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware.
By Yuan Zeng, Yujia Shi, Tiao Tan, Xingting Li, Yaqi Qin, Zongqing Lu, Wenming Yang, Jing-Hao Xue, Qingmin Liao
arXiv:2606. 08107v1 Announce Type: cross Abstract: Robotics faces a fundamental challenge of data scarcity.
By Ji Woong Kim, Ke Wang, Zipeng Fu, Sirui Chen, Cong Zhao, Jeff Lai, Chelsea Finn
The paper introduces 3DROID, a dataset of renderable 3D Gaussian scenes that are anchored to a robot’s metric workspace and include per-scene reliability metrics. It examines how the reliability of camera extrinsics and pose conditioning affect the fidelity of 3D Gaussian representations, proposing a calibration-aware pipeline that improves novel-view rendering when extrinsics are trustworthy. The resulting dataset, available on Hugging Face, provides robot manipulation researchers with metric-scale, pose-anchored 3D data and reliability annotations.
By Wonguen Cho, Junhoo Lee, Nojun Kwak
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable.
arXiv:2608.29927v1 Announce Type: new
Abstract: We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each indivi...
By Daeyun Shin, Yunhan Zhao, Shu Kong, Alexander C. Berg, Charless Fowlkes
MEgoVista is an offline pipeline that converts a single unprepared egocentric video into metric two‑hand and head motion within a gravity‑aligned world frame. It uniquely reconstructs motion in environments beyond studio volumes, uses calibrated stereo for absolute scale, and evaluates its outputs against independent optical capture to audit accuracy. The system thus expands the settings where high‑fidelity hand‑motion labels can be generated from natural, head‑worn recordings.
By Jiangong Xiao (Northwestern Polytechnical University), Zhihao Zhang (Xi'an Jiaotong University), Yifei Dong (Maniformer), Chao Ma (Maniformer), Zhouyi Jin (Maniformer), Zhiwen Hou (Maniformer), Li Liu (Maniformer), Weihuang Chen (Xi'an Jiaotong University), Hongbin Sun (Xi'an Jiaotong University), Maoqing Yao (Maniformer)
arXiv:2607. 17790v1 Announce Type: cross Abstract: Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment.
By Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang