arXiv Machine Learning By Junfeng Long, Pieter Abbeel, Koushil Sreenath, Roberto Horowitz, Guanya Shi, C. Karen Liu

Humanoid Safe Stop via Learned Stoppability Value

Read the original on arXiv Machine Learning →

The paper introduces Safe-Stop, a task‑agnostic framework for humanoid robots that determines whether an emergency stop command can be safely executed from the current state. It couples a learned stop policy with two complementary stoppability estimators: a stop‑probability estimator trained on outcomes of a fixed stop policy, and a reach‑avoidance estimator trained via Hamilton‑Jacobi backup. By requiring agreement between both estimators before committing to a stop, Safe‑Stop can robustly decide to stop or hand off to a fall‑damping fallback without needing retraining for different upstream tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv Machine Learning
Sep 24

ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control

ForgetMimic is a motion-level unlearning method for reinforcement learning-based humanoid control. It selectively degrades performance on a chosen subset of motions while preserving the policy’s effectiveness on the remaining motions. Experiments on Unitree G1 and H2 robots across 12 motions show that the method successfully removes memory of designated motions without affecting other behaviors.

By Xukun Luan, Zhongxiang Lei, Chen Gong, Shaowei Li, Yuanguo Bi, Jinyan Liu