Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulation. We present SiMDex, a similarity-based data mining framework that casts human data selection for VLA post-training in dexterous manipulation as a recommendation problem.
arXiv:2609.13851v1 Announce Type: cross
Abstract: Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data...
By Chenwei Wang, Dianye Huang, Match W. L. Ko, Chenjia Bai, Zhongliang Jiang
RoboTok is an internet‑scale data engine that retrieves human manipulation videos from the web to train dexterous robot policies. It learns a latent motion space from 3D hand trajectories in actor‑centered reference frames, allowing manipulation behaviors to be compared across different viewpoints, scenes, and occlusions while remaining compact for efficient search. Experiments show RoboTok retrieves more relevant demonstrations and improves downstream robot task success compared to existing retrieval methods.
By Howard Qian, Yiting Chen, Yunfei Xie, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen, Chen Wei, Kaiyu Hang
arXiv:2606. 11628v1 Announce Type: cross Abstract: The most widely-adopted robot learning pipelines today learn skills from robot demonstrations or structured human data, which are expensive to collect and tied to specific embodiments.
By Harsh Gupta, Guanya Shi, Wenzhen Yuan
Ego4WAM investigates how various properties of egocentric human data—such as human‑robot alignment, data duration, task diversity, and supervision type—affect robot learning. The study, conducted under a unified world‑action model framework, shows that aligned demonstrations improve out‑of‑distribution generalization and lower the amount of target‑task robot data needed. It also finds that video‑only supervision remains effective, and that data duration and task diversity influence downstream capabilities in distinct ways, as validated on real robots and RoboDojo.
By Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu
Sim-and-Human Co-training (SimHum) is a method that combines simulation and human demonstration data to train bimanual manipulation policies. It first extracts kinematic priors from simulation and visual priors from human observations, then fine‑tunes on a small real‑robot dataset. With only 80 real‑robot episodes per task, SimHum achieves 62.5% success on out‑of‑distribution scenes across four tabletop tasks, outperforming real‑only training by 53.7% and improving the best single‑source baseline by 35.0% in a matched‑time study.
By Kaipeng Fang, Weiqing Liang, Yuyang Li, Ji Zhang, Pengpeng Zeng, Heng Tao Shen, Jingkuan Song, Lianli Gao
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity.
arXiv:2609.36691v1 Announce Type: new
Abstract: Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embod...
By Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu, Haoran Lu, Yijiang Li, Yaqi Xie, Katia P. Sycara, Han Liu
HumanEgo is a framework that enables zero‑shot robot learning from short egocentric human videos by converting each demonstration into an entity‑level hand‑object interaction representation and training a flow‑matching policy with dense auxiliary objectives. The method is robot‑data‑free, hardware‑agnostic, and data‑efficient, achieving 92.5 % success on four real‑world tasks with only 30 minutes of human video per task and outperforming matched‑time robot teleoperation by 41 %. HumanEgo also robustly transfers zero‑shot across new robots, cameras, and environments, and is released as an open‑source tool for learning robot policies directly from human data.
By Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos
arXiv:2602. 13197v2 Announce Type: replace-cross Abstract: The ability to learn manipulation skills by watching videos of humans has the potential to unlock a new source of highly scalable data for robot learning.
By Albert J. Zhai, Kuo-Hao Zeng, Jiasen Lu, Ali Farhadi, Shenlong Wang, Wei-Chiu Ma
arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.
By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin