arXiv:2606. 06493v1 Announce Type: cross Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.
By Lizhi Yang, Junheng Li, Nehar Poddar, Yiling Hou, Gio Huh, Robert Griffin, Georgia Gkioxari, Aaron Ames
ULTRA is a unified framework for autonomous humanoid whole-body locomotion and manipulation that overcomes limitations of prior methods by combining a physics-driven neural retargeting algorithm with a multimodal controller. The retargeting algorithm translates large-scale motion capture data into physically plausible humanoid motions, while the controller learns to handle both dense motion references and sparse task specifications using a range of sensory inputs, from accurate motion-capture states to noisy egocentric vision. In simulation and on a real Unitree G1 humanoid, ULTRA demonstrates improved generalization and robustness, enabling coordinated whole-body behavior from sparse intent without relying on test-time reference motions.
By Xialin He, Sirui Xu, Xinyao Li, Runpei Dong, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
The paper introduces InterTrack, a behavior world model that enables humanoid robots to perform robust whole-body tracking while interacting with varied terrain and objects. Using a Transformer architecture, InterTrack predicts actions, states, and behavior distributions conditioned on the environment, and it is trained with an automated pipeline that reconstructs 3D support geometry from retargeted motions. The system achieves an 81.3% success rate on terrain interaction, a 99.3% fall-recovery rate, and outperforms leading baselines in both free-space tracking and cross-terrain scenarios.
By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Borui Zhang, Bingyao Yu, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv:2606. 14218v1 Announce Type: cross Abstract: For robots to work safely in household environments, they need to be compliant and react to torque and force feedback during contact.
By Litian Liang, Jingxi Xu, Xinda Qi, Yujun Cai, Houzhu Ding, Luqi Wang, Zhixin Sun, Jyh-Herng Chow, Ming Yang, Mark Cutkosky
FWBC‑VLA is a force‑aware framework that links vision‑language‑action (VLA) models with whole‑body compensation control for wheeled‑legged robots. It introduces HSR‑Force, a sensorless residual‑torque estimator that infers contact strength and encodes this information as tokens for the VLA action decoder, allowing the policy to detect contact onset, loading, and release. The system fine‑tunes a pretrained VLA backbone on a large WL&Arm dataset, combines proprioceptive, Jacobian‑derived force, and contact estimates to generate corrective actions, and demonstrates effectiveness in real‑world tasks such as whiteboard wiping and door opening.
By Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong We, Qiaojun Yu, Dibo Hou
arXiv:2609.16683v1 Announce Type: cross
Abstract: Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and objec...
By Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu, Mingzhi Pei, Ruoqu Chen, Mengdi Xu