arXiv AI By Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

Read the original on arXiv AI →

arXiv:2608. 16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 14

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

arXiv:2511. 07820v4 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown for humanoid control.

By Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Fernando Casta\~neda, Sirui Chen, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Jinhyung Park, David Sami, Zi Wang, Xingye Da, Runyu Ding, Cyrus Hogg, Lina Song, Edy Lim, Eugene Jeong, Tairan He, Haoru Xue, Wenli Xiao, Simon Yuen, Jan Kautz, Yan Chang, Umar Iqbal, Linxi "Jim" Fan, Yuke Zhu
arXiv AI
Jul 17

Scaling Behavior Foundation Model for Humanoid Robots

arXiv:2607. 15163v1 Announce Type: cross Abstract: Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents.

By Weishuai Zeng, Kangning Yin, Xiaojie Niu, Shunlin Lu, Weixiang Zhong, Jiahe Chen, Feiyu Jia, Xiao Chen, Zirui Wang, Furui Xu, Ming Zhou, Kailin Li, Weinan Zhang, He Wang, Li Yi, Dahua Lin, Jiangmiao Pang, Jingbo Wang
arXiv Computer Vision
Aug 27

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming proposes PoseOFF, a representation that captures local motion around human joints by conditioning optical flow extraction on pose. This structured motion representation aligns with human kinematics and improves early action recognition accuracy across multiple datasets and backbones. PoseOFF achieves comparable or better performance while observing less of the action sequence, making it suitable for real‑time, resource‑constrained robotic systems.

By Lewis de Zoete Grundy, Chris McCarthy, Christopher Fluke
arXiv Computer Vision
Sep 25

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

BeyondRetarget is an end‑to‑end framework that learns to generate executable humanoid robot motions directly from monocular RGB videos, bypassing the need for an explicit human motion representation. By learning robot‑oriented implicit representations and incorporating a contact‑aware motion optimization mechanism, the method captures cross‑morphology motion structures and improves temporal consistency and physical plausibility. Experiments demonstrate that BeyondRetarget achieves higher execution success rates, lower latency, and greater accuracy and robustness in both simulation and real humanoid robots.

By Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao