HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers
arXiv:2606. 06493v1 Announce Type: cross Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.
arXiv:2603. 03751v2 Announce Type: replace-cross Abstract: Cooperative object transport in unstructured environments remains challenging for assistive humanoids because strong, time-varying interaction forces can make tracking-centric whole-body control unreliable, especially in close-contact support tasks.
arXiv:2606. 06493v1 Announce Type: cross Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.
ULTRA is a unified framework for autonomous humanoid whole-body locomotion and manipulation that overcomes limitations of prior methods by combining a physics-driven neural retargeting algorithm with a multimodal controller. The retargeting algorithm translates large-scale motion capture data into physically plausible humanoid motions, while the controller learns to handle both dense motion references and sparse task specifications using a range of sensory inputs, from accurate motion-capture states to noisy egocentric vision. In simulation and on a real Unitree G1 humanoid, ULTRA demonstrates improved generalization and robustness, enabling coordinated whole-body behavior from sparse intent without relying on test-time reference motions.
The paper introduces InterTrack, a behavior world model that enables humanoid robots to perform robust whole-body tracking while interacting with varied terrain and objects. Using a Transformer architecture, InterTrack predicts actions, states, and behavior distributions conditioned on the environment, and it is trained with an automated pipeline that reconstructs 3D support geometry from retargeted motions. The system achieves an 81.3% success rate on terrain interaction, a 99.3% fall-recovery rate, and outperforms leading baselines in both free-space tracking and cross-terrain scenarios.
arXiv:2606. 14218v1 Announce Type: cross Abstract: For robots to work safely in household environments, they need to be compliant and react to torque and force feedback during contact.
FWBC‑VLA is a force‑aware framework that links vision‑language‑action (VLA) models with whole‑body compensation control for wheeled‑legged robots. It introduces HSR‑Force, a sensorless residual‑torque estimator that infers contact strength and encodes this information as tokens for the VLA action decoder, allowing the policy to detect contact onset, loading, and release. The system fine‑tunes a pretrained VLA backbone on a large WL&Arm dataset, combines proprioceptive, Jacobian‑derived force, and contact estimates to generate corrective actions, and demonstrates effectiveness in real‑world tasks such as whiteboard wiping and door opening.
arXiv:2609.16683v1 Announce Type: cross Abstract: Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and objec...
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of...
arXiv:2605. 23733v2 Announce Type: replace-cross Abstract: Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate diverse motions with high fidelity.
GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.
Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipulate it, and then resuming locomotion. It also commonly relies on low degree-of-freedom (DoF) end effectors that behave like an open-close grasp primitive.
arXiv:2603.03768v2 Announce Type: replace-cross Abstract: Full-stack human-robot collaboration (HRC) can become brittle when replacing a planner, partner model, coordination policy, or controller cha...
arXiv:2606. 29209v1 Announce Type: cross Abstract: We present AnyBody, a unified whole-body humanoid controller driven by an arbitrary subset of body keypoints chosen at deploy time.