HumanEgo is a framework that enables zero‑shot robot learning from short egocentric human videos by converting each demonstration into an entity‑level hand‑object interaction representation and training a flow‑matching policy with dense auxiliary objectives. The method is robot‑data‑free, hardware‑agnostic, and data‑efficient, achieving 92.5 % success on four real‑world tasks with only 30 minutes of human video per task and outperforming matched‑time robot teleoperation by 41 %. HumanEgo also robustly transfers zero‑shot across new robots, cameras, and environments, and is released as an open‑source tool for learning robot policies directly from human data.
By Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos
arXiv:2606. 30645v1 Announce Type: cross Abstract: Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion.
By Yen-Jen Wang, Jiaman Li, Sirui Chen, Takara E. Truong, Pei Xu, Pieter Abbeel, Rocky Duan, Koushil Sreenath, Angjoo Kanazawa, Carmelo Sferrazza, Guanya Shi, Karen Liu
ULTRA is a unified framework for autonomous humanoid whole-body locomotion and manipulation that overcomes limitations of prior methods by combining a physics-driven neural retargeting algorithm with a multimodal controller. The retargeting algorithm translates large-scale motion capture data into physically plausible humanoid motions, while the controller learns to handle both dense motion references and sparse task specifications using a range of sensory inputs, from accurate motion-capture states to noisy egocentric vision. In simulation and on a real Unitree G1 humanoid, ULTRA demonstrates improved generalization and robustness, enabling coordinated whole-body behavior from sparse intent without relying on test-time reference motions.
By Xialin He, Sirui Xu, Xinyao Li, Runpei Dong, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
arXiv:2608. 14028v1 Announce Type: cross Abstract: Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments.
By Zhiyue Zhao, Jingyi Wu, Hairuo Liu, Mingyu Liu, Liyang Li, Hengdi Zhang, Tong He, Zhengxue Cheng
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale.
arXiv:2603. 09170v2 Announce Type: replace-cross Abstract: Achieving versatile and natural whole-body humanoid interaction control remains challenging due to the high cost of whole-body teleoperation data.
By Haoran Yang, Jiacheng Bao, Yucheng Xin, Haoming Song, Yuyang Tian, Bin Zhao, Dong Wang, Xuelong Li
AtomEgo investigates how to integrate large-scale egocentric human interaction data into embodied foundation model pre‑training. The study uses a curated 2,659‑hour corpus and a scalable data pipeline to evaluate three co‑training paradigms across vision‑language‑action and world‑action architectures. Results show that the benefit of egocentric data depends on both its scale and the quality of alignment with robotic embodiment, offering practical guidance for scalable ego‑robot pre‑training.
By Di Wu, Dongchen Zheng, Junhe Sheng, Zhongxing Wei, Songxin Zhang, Zejian Xie, Xiaoquan Sun, Junyang Zheng, Zhuoyang Song, Jiaxing Zhang, Jiayu Chen
arXiv:2606. 28813v1 Announce Type: cross Abstract: Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactions.
By Shuo Cheng, Chuye Zhang, Alfred Cueva, Caelan Garrett, Ajay Mandlekar, Danfei Xu
arXiv:2606. 17846v1 Announce Type: cross Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale.
By Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen
arXiv:2609.22332v1 Announce Type: cross
Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act...
By Jiadi You, Qize Yu, Yue Chen, Minghong Cai, Zhide Zhong, Yuran Wang, Bowen Ping, Jiaqi Liang, Zhenhao Shen, Haodong Yan, Yinchuan Li, Ruihai Wu, Xiaojuan Qi, Yingcong Chen
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively.
arXiv:2608. 16837v1 Announce Type: cross Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation.
By Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang