The paper introduces InterTrack, a behavior world model that enables humanoid robots to perform robust whole-body tracking while interacting with varied terrain and objects. Using a Transformer architecture, InterTrack predicts actions, states, and behavior distributions conditioned on the environment, and it is trained with an automated pipeline that reconstructs 3D support geometry from retargeted motions. The system achieves an 81.3% success rate on terrain interaction, a 99.3% fall-recovery rate, and outperforms leading baselines in both free-space tracking and cross-terrain scenarios.
By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Borui Zhang, Bingyao Yu, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv:2606. 12814v1 Announce Type: cross Abstract: Recent reinforcement learning approaches have shown great promise in improving humanoid motion tracking performance and achieving fall recovery under disturbances.
By Xiao Ren, Yuhui Yang, Zongbiao Weng, Zhijie Liu, He Kong
arXiv:2606. 08253v1 Announce Type: cross Abstract: Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate robustly, safely, and accurately.
By Alessandro Montenegro, Shihao Li, Puze Liu, Alberto Maria Metelli, Jan Peters
arXiv:2606. 06493v1 Announce Type: cross Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.
By Lizhi Yang, Junheng Li, Nehar Poddar, Yiling Hou, Gio Huh, Robert Griffin, Georgia Gkioxari, Aaron Ames
arXiv:2609.38653v1 Announce Type: cross
Abstract: Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human mo...
By Merkourios Simos, Chengkun Li, Bianca Ziliotto, Alexander Mathis
Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipulate it, and then resuming locomotion. It also commonly relies on low degree-of-freedom (DoF) end effectors that behave like an open-close grasp primitive.