Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs.
arXiv:2608. 16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions.
By Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han
arXiv:2507. 19684v2 Announce Type: replace-cross Abstract: Socially interactive humanoid robots must engage with humans through their bodies, adapting in real time to a partner's movement, intent, and abilities.
By Bermet Burkanova, Yasaman Etesam, Payam Jome Yazdian, Trinity Evans, Chuxuan Zhang, Zoe Stanley, Paige Tutt\"os\'i, Angelica Lim
DirtyMoCap is a marker‑layout‑free framework that converts unordered, noisy optical motion capture markers into a fixed set of proxy anchors representing skeletal joints and body surface points. Using a recurrent sliding‑window architecture to track these anchors and a custom differentiable Gauss‑Newton solver to fit the SMPL‑H model, the method learns adaptive observation confidence, smoothness, and prior weights end‑to‑end. Experiments show that DirtyMoCap generalizes across arbitrary marker configurations, outperforms configuration‑specific baselines in joint and vertex accuracy, and achieves up to a 100× speedup over standard PyTorch implementations, enabling the creation of a temporally coherent Kung Fu motion dataset.
By Long Wang, Shuting Zhao, Shen Yan, Siyuan Yu, Xiaoben Li, Zeyu Cai, Yumeng Hou, Yuliang Xiu
Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming proposes PoseOFF, a representation that captures local motion around human joints by conditioning optical flow extraction on pose. This structured motion representation aligns with human kinematics and improves early action recognition accuracy across multiple datasets and backbones. PoseOFF achieves comparable or better performance while observing less of the action sequence, making it suitable for real‑time, resource‑constrained robotic systems.
By Lewis de Zoete Grundy, Chris McCarthy, Christopher Fluke
arXiv:2606. 06627v1 Announce Type: cross Abstract: Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resemble robot behavior and 3D hand poses are captured with specialized hardware.
By Richard Li, Aditya Prakash, Andrew Wen, Saurabh Gupta, Yilun Du, Pulkit Agrawal
Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally. However, many existing approaches to human act...
arXiv:2606. 09842v1 Announce Type: cross Abstract: Applying Human Pose Estimation (HPE) in real world environments remains a challenging task, this paper explores and surveys real time HPE approaches and their limitations in sports analysis for individuals, alongside developing a practical lightweight prototype for real world testing and usage.
By Parth Agrawal, Ronit, Sagar Kumar, Aashish Bhambri
arXiv:2609.36628v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, partic...
By Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe, Zhaohe Liao
arXiv:2608. 19480v1 Announce Type: new Abstract: Human pose estimation has advanced significantly due to the development of deep learning models, increased data availability, and improved computing resources.
By Luis F. Gomez, Julian Fierrez, Roberto Daza, Ruben Tolosana, Aythami Morales, Gonzalo Garrido, Javier Rueda, Enrique Navarro
arXiv:2609.37181v1 Announce Type: cross
Abstract: Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has e...
By Jin Chen, Yiming Jiang, Chongyang Xu, Modi Shi, Shijia Peng, Li Chen, Tianyu Li, Mu Xu, Yilun Chen, Steven Hoi, Hongyang Li
BeyondRetarget is an end‑to‑end framework that learns to generate executable humanoid robot motions directly from monocular RGB videos, bypassing the need for an explicit human motion representation. By learning robot‑oriented implicit representations and incorporating a contact‑aware motion optimization mechanism, the method captures cross‑morphology motion structures and improves temporal consistency and physical plausibility. Experiments demonstrate that BeyondRetarget achieves higher execution success rates, lower latency, and greater accuracy and robustness in both simulation and real humanoid robots.
By Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao