arXiv Machine Learning

PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control

arXiv Machine Learning
Jul 31

REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning

arXiv:2603. 13707v3 Announce Type: replace-cross Abstract: Humanoid loco-manipulation requires coordinated task-space motion planning with stable loco-manipulation command tracking under complex robot-environment dynamics and long-horizon tasks.

By Zhaoyuan Gu, Yipu Chen, Zimeng Chai, Alfred Cueva, Thong Nguyen, Yifan Wu, Huishu Xue, Minji Kim, Isaac Legene, Fukang Liu, KyoungMok Kim, Ayan Barula, Yongxin Chen, Ye Zhao
arXiv Machine Learning
Sep 14

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.

By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
arXiv AI
Aug 20

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.

By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv Computer Vision
Sep 18

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.

By Zhongbo Zhang, Zaibin Zhang, Yifan Wang, Changbo Yan, Lijun Wang, Huchuan Lu
arXiv Machine Learning
Sep 7

SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control

SCRIPT is a scalable diffusion policy that uses a Joint Action-State-Text Diffusion Transformer (JAST‑DiT) to jointly encode actions, physical states, and natural‑language instructions, enabling direct interaction between language semantics and control dynamics. The method employs a multi‑stage training framework, including supervised imitation pre‑training, a nonlinear history conditioning mechanism for stable autoregressive control, and a post‑training stage with Reinforcement Learning with Hybrid Rewards (RLHR) that injects learnable noise to improve motion quality and instruction following. Experiments on the 1200‑hour MotionMillion dataset show that SCRIPT outperforms prior state‑of‑the‑art methods across text alignment, motion quality, and physical realism, and its performance scales consistently with model size.

By Jingyan Zhang, Han Liang, Ruichi Zhang, Bin Li, Juze Zhang, Xin Chen, Jingya Wang, Lan Xu, Jingyi Yu
arXiv AI
Jun 30

ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control

arXiv:2606. 30362v1 Announce Type: cross Abstract: While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motions.

By Xiao Chen, Weishuai Zeng, Xiaojie Niu, Zirui Wang, Jianan Li, Huayi Wang, Furui Xu, Jiahe Chen, Weixiang Zhong, Lihe Ding, Kailin Li, Jiangmiao Pang, Tai Wang, Tianfan Xue, Jingbo Wang
arXiv Computer Vision
Sep 21

ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

ULTRA is a unified framework for autonomous humanoid whole-body locomotion and manipulation that overcomes limitations of prior methods by combining a physics-driven neural retargeting algorithm with a multimodal controller. The retargeting algorithm translates large-scale motion capture data into physically plausible humanoid motions, while the controller learns to handle both dense motion references and sparse task specifications using a range of sensory inputs, from accurate motion-capture states to noisy egocentric vision. In simulation and on a real Unitree G1 humanoid, ULTRA demonstrates improved generalization and robustness, enabling coordinated whole-body behavior from sparse intent without relying on test-time reference motions.

By Xialin He, Sirui Xu, Xinyao Li, Runpei Dong, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
arXiv Machine Learning
Aug 27

LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

LM‑X is a generalist vision‑language‑action policy that augments action prediction with three online, explicitly supervised signals: return‑to‑go (RTG) for task progress, event‑to‑go (ETG) for the next semantic transition, and heteroscedastic action flow for local reliability. By conditioning action generation on these signals, LM‑X embeds explainability directly into control rather than as a post‑hoc explanation. After a 20‑day pretraining run on 64 GPUs, LM‑X outperforms an action‑only backbone by 16.0 points and a single‑head variant by 10.8 points, and achieves 74.1 % success on 50 RoboTwin2.0 tasks and 68.6 % on seven real‑robot tasks, surpassing the GR00T N1.7 baseline.

By Jin Lou, Jingxuan Zhu, Andong Chen, Xupeng Wang, Yuan Xu, Yuexuan Li, Xingdong Zhu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Jingyi Li, Liangliang Chen, Jinyan Liu, Zhiqi Song, Jidong Zhang, Hongming Li, Yuchen Zhu
arXiv Machine Learning
Jul 28

Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation

arXiv:2607. 24083v1 Announce Type: new Abstract: Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process.

By Valerio Belli (UNIROMA, UCL), Valerio Modugno (UCL), Enrico Mingo Hoffman (HUCEBOT), Fabio Amadio (HUCEBOT)
arXiv AI
Sep 16

SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating

SafeFlow is a real‑time, text‑driven humanoid control framework that blends physics‑guided motion generation with a three‑stage safety gate. It uses Physics‑Guided Rectified Flow Matching in a VAE latent space to produce physically executable trajectories, accelerates sampling with Reflow, and filters unsafe outputs via semantic OOD detection, directional sensitivity checks, and hard kinematic constraints before handing them to a motion‑tracking controller. Experiments on the Unitree G1 show that SafeFlow achieves higher success rates, better physical compliance, and faster inference than diffusion‑ and retargeting‑based baselines while maintaining motion diversity.

By Hanbyel Cho, Sang-Hun Kim, Jeonguk Kang, Donghan Koo