UniExo is a framework that builds a single, multi-skill musculoskeletal human policy by distilling four imitation experts—walking, turning, running, and backward walking—into one network guided by a skill latent. The human policy is fine‑tuned with reinforcement learning on transition sequences, achieving a 94.7% tracking success rate on unseen clips and greater robustness to perturbations. A hip exoskeleton controller is then co‑adapted with this human policy via multi‑agent reinforcement learning, enabling it to assist across four treadmill speeds and a continuous route of all four skills without explicit mode switching.
By Yifei Yuan, Jakob Wolf, Ghaith Androwis, Xianlian Zhou
arXiv:2607. 12114v1 Announce Type: cross Abstract: A humanoid that can walk should not relearn locomotion from scratch to jog or run.
By Kwan-Yee Lin, Zilin Wang, Janelle J. Liu, Stella X. Yu
arXiv:2606. 15896v1 Announce Type: cross Abstract: Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective.
By Loukas Kordos, Leonard T. Franz, Simon Rappenecker, Oliver Hausdoerfer, Angela P. Schoellig, Pavel Kolev, Georg Martius
arXiv:2607. 29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings.
By Yuxin Chen, Hari Srikanth, Nathan Jew, Menglin Wu, Pengcheng Wang, Junli Ren, Masayoshi Tomizuka, Peng Xu, Jinyu Xie, Thomas Tian
The paper investigates how reinforcement‑learning policies for legged robots encode gait information by examining the effective rank of the policy Jacobian conditioned on gait phase. It finds that common architectural features such as layer normalization and residual connections allocate more representational capacity to swing than stance, a pattern absent in vanilla MLPs. Leveraging these insights, the authors propose a simple recipe that improves sim‑to‑real transfer, reducing joint jitter on a physical Spot robot by roughly three‑fold.
By Felipe Tommaselli, Thiago H. Segreto, Juliano D. Negri, Ricardo V. Godoy, Marcelo Becker
arXiv:2608. 02069v1 Announce Type: cross Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training times.
By Martin Opat
arXiv:2607. 24083v1 Announce Type: new Abstract: Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process.
By Valerio Belli (UNIROMA, UCL), Valerio Modugno (UCL), Enrico Mingo Hoffman (HUCEBOT), Fabio Amadio (HUCEBOT)
arXiv:2606. 04718v1 Announce Type: cross Abstract: Humans primarily rely on walking and running to traverse complex terrains, without resorting to unnecessarily complex motion patterns.
By Kailun Huang (Hong Kong University of Science and Technology), Zikang Xie (Hong Kong University of Science and Technology), Yanzhe Xie (Hong Kong University of Science and Technology), Panpan Liao (Guangdong University of Technology), Fanghai Zhang (Hong Kong University of Science and Technology), Yanheng Mai (Hong Kong University of Science and Technology), Wenhao Xu (South China Agricultural University), Yunheng Wang (Hong Kong University of Science and Technology), Renjing Xu (Hong Kong University of Science and Technology), Haohui Huang (Guangdong University of Technology)
arXiv:2607. 09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments.
By Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu
GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.
By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv:2608. 00715v1 Announce Type: cross Abstract: Learning-based controllers can deliver exoskeleton assistance after training entirely in physics-based simulation, yet few controllers that address human-device co-adaptation have been validated on real users by whole-body metabolic measurement, the standard benchmark for assistive walking.
By Yifei Yuan, Jakob Wolf, Ghaith Androwis, Xianlian Zhou
arXiv:2608. 01506v1 Announce Type: cross Abstract: Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies often break when hardware properties shift.
By Dichen Li, Bo Ai, Nico Bohlinger, Jan Peters, Hao Su, Henrik I. Christensen