arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.
By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv:2606. 12814v1 Announce Type: cross Abstract: Recent reinforcement learning approaches have shown great promise in improving humanoid motion tracking performance and achieving fall recovery under disturbances.
By Xiao Ren, Yuhui Yang, Zongbiao Weng, Zhijie Liu, He Kong
arXiv:2606. 17011v1 Announce Type: cross Abstract: Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models.
By Wei Xiao, Weiliang Tang, Yuying Ge, Hui Zhou, Yao Mu, Li Zhang, Yixiao Ge
arXiv:2603. 15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints.
By Mumuksh Tayal, Manan Tayal, Ravi Prakash
ForgetMimic is a motion-level unlearning method for reinforcement learning-based humanoid control. It selectively degrades performance on a chosen subset of motions while preserving the policy’s effectiveness on the remaining motions. Experiments on Unitree G1 and H2 robots across 12 motions show that the method successfully removes memory of designated motions without affecting other behaviors.
By Xukun Luan, Zhongxiang Lei, Chen Gong, Shaowei Li, Yuanguo Bi, Jinyan Liu
GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.
By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
The paper introduces a two‑step deep reinforcement learning framework for Reach‑Avoid‑Stay (RAS) problems, aiming to compute the maximal robust RAS set and its control policy for general dynamic systems. First, it learns the maximal robust control‑invariant set inside the target and a policy to keep the system within it. Then it uses this invariant set as a target to compute the maximal robust reach‑avoid set, proving equivalence to the maximal robust RAS set and constructing a switching policy that guarantees task completion. Simulation results show the method achieves exact maximal RAS sets without training errors and outperforms baseline approaches in accuracy and performance.
By Gabriel Chenevert, Jingqi Li, Achyuta kannan, Sangjae Bae, Donggun Lee
arXiv:2609. 12014v1 Announce Type: new Abstract: Safe offline reinforcement learning assumes a cost function on every transition.
By Adam Haroon, Cody Fleming
arXiv:2602.04599v3 Announce Type: replace
Abstract: We propose stochastic decision horizons (SDH), a theoretically grounded framework for survival-based constrained reinforcement learning from per-st...
By Nikola Milosevic, Leonard Franz, Daniel Haeufle, Georg Martius, Nico Scherf, Pavel Kolev
arXiv:2508. 16943v3 Announce Type: replace-cross Abstract: Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism.
By Haozhuo Zhang, Jingkai Sun, Michele Caprio, Angelo Cangelosi, Jian Tang, Shanghang Zhang, Qiang Zhang, Wei Pan
arXiv:2606. 14585v1 Announce Type: cross Abstract: Generative dynamics models enable planning in challenging robotic systems, but safe deployment requires reliably detecting policy-induced out-of-distribution (OOD) transitions.
By Hongzhan Yu, Chenghao Li, Ruipeng Zhang, Henrik Christensen, Sicun Gao