arXiv Computer Vision

RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion

arXiv Computer Vision
Sep 25

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

BeyondRetarget is an end‑to‑end framework that learns to generate executable humanoid robot motions directly from monocular RGB videos, bypassing the need for an explicit human motion representation. By learning robot‑oriented implicit representations and incorporating a contact‑aware motion optimization mechanism, the method captures cross‑morphology motion structures and improves temporal consistency and physical plausibility. Experiments demonstrate that BeyondRetarget achieves higher execution success rates, lower latency, and greater accuracy and robustness in both simulation and real humanoid robots.

By Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao
arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
Hugging Face Trending Papers
Jun 18

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity.

arXiv AI
Jul 1

A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

arXiv:2509. 15443v2 Announce Type: replace-cross Abstract: Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing abundant, large-scale human motion collections.

By Xingyu Chen, Hanyu Wu, Sikai Wu, Mingliang Zhou, Diyun Xiang, Haodong Zhang, Yangchen Zhou, Yukang Gao, Yi Gu, Renjing Xu
arXiv Computer Vision
Sep 17

StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions

arXiv:2609.18430v1 Announce Type: new Abstract: Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucP...

By WM Team, Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang, Kaicheng Yu