UniExo is a framework that builds a single, multi-skill musculoskeletal human policy by distilling four imitation experts—walking, turning, running, and backward walking—into one network guided by a skill latent. The human policy is fine‑tuned with reinforcement learning on transition sequences, achieving a 94.7% tracking success rate on unseen clips and greater robustness to perturbations. A hip exoskeleton controller is then co‑adapted with this human policy via multi‑agent reinforcement learning, enabling it to assist across four treadmill speeds and a continuous route of all four skills without explicit mode switching.
By Yifei Yuan, Jakob Wolf, Ghaith Androwis, Xianlian Zhou
arXiv:2607. 12114v1 Announce Type: cross Abstract: A humanoid that can walk should not relearn locomotion from scratch to jog or run.
By Kwan-Yee Lin, Zilin Wang, Janelle J. Liu, Stella X. Yu
arXiv:2606. 15896v1 Announce Type: cross Abstract: Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective.
By Loukas Kordos, Leonard T. Franz, Simon Rappenecker, Oliver Hausdoerfer, Angela P. Schoellig, Pavel Kolev, Georg Martius
arXiv:2607. 29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings.
By Yuxin Chen, Hari Srikanth, Nathan Jew, Menglin Wu, Pengcheng Wang, Junli Ren, Masayoshi Tomizuka, Peng Xu, Jinyu Xie, Thomas Tian
The paper investigates how reinforcement‑learning policies for legged robots encode gait information by examining the effective rank of the policy Jacobian conditioned on gait phase. It finds that common architectural features such as layer normalization and residual connections allocate more representational capacity to swing than stance, a pattern absent in vanilla MLPs. Leveraging these insights, the authors propose a simple recipe that improves sim‑to‑real transfer, reducing joint jitter on a physical Spot robot by roughly three‑fold.
By Felipe Tommaselli, Thiago H. Segreto, Juliano D. Negri, Ricardo V. Godoy, Marcelo Becker
arXiv:2608. 02069v1 Announce Type: cross Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training times.
By Martin Opat