arXiv Machine Learning

Humanoid Safe Stop via Learned Stoppability Value

The paper introduces Safe-Stop, a task‑agnostic framework for humanoid robots that determines whether an emergency stop command can be safely executed from the current state. It couples a learned stop policy with two complementary stoppability estimators: a stop‑probability estimator trained on outcomes of a fixed stop policy, and a reach‑avoidance estimator trained via Hamilton‑Jacobi backup. By requiring agreement between both estimators before committing to a stop, Safe‑Stop can robustly decide to stop or hand off to a fall‑damping fallback without needing retraining for different upstream tasks.

arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv Machine Learning
Sep 24

ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control

ForgetMimic is a motion-level unlearning method for reinforcement learning-based humanoid control. It selectively degrades performance on a chosen subset of motions while preserving the policy’s effectiveness on the remaining motions. Experiments on Unitree G1 and H2 robots across 12 motions show that the method successfully removes memory of designated motions without affecting other behaviors.

By Xukun Luan, Zhongxiang Lei, Chen Gong, Shaowei Li, Yuanguo Bi, Jinyan Liu
arXiv AI
Aug 20

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.

By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv Machine Learning
Sep 3

Deep Reinforcement Learning for Reach-Avoid-Stay Problems

The paper introduces a two‑step deep reinforcement learning framework for Reach‑Avoid‑Stay (RAS) problems, aiming to compute the maximal robust RAS set and its control policy for general dynamic systems. First, it learns the maximal robust control‑invariant set inside the target and a policy to keep the system within it. Then it uses this invariant set as a target to compute the maximal robust reach‑avoid set, proving equivalence to the maximal robust RAS set and constructing a switching policy that guarantees task completion. Simulation results show the method achieves exact maximal RAS sets without training errors and outperforms baseline approaches in accuracy and performance.

By Gabriel Chenevert, Jingqi Li, Achyuta kannan, Sangjae Bae, Donggun Lee
arXiv AI
Jun 15

Sensitivity Shaping for Latent Modeling

arXiv:2606. 14585v1 Announce Type: cross Abstract: Generative dynamics models enable planning in challenging robotic systems, but safe deployment requires reliably detecting policy-induced out-of-distribution (OOD) transitions.

By Hongzhan Yu, Chenghao Li, Ruipeng Zhang, Henrik Christensen, Sicun Gao