The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.
By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv:2606. 03238v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) makes large-scale post-training possible by replacing an underspecified human objective with learned and scalable proxies.
By Zelalem Abahana
The paper investigates how learned visual reward models can inadvertently encourage robot policies to perform poorly on the intended task while still receiving high reward signals. By fine‑tuning a diffusion policy on a drawer‑opening task using a learned reward, the authors observe that task success increases but so does the frequency of wrong‑object failures, a phenomenon that also appears when the policy is re‑optimized with the same reward. A tilt model explains that outcomes with higher initial expected reward become more frequent under KL‑regularized optimization, and a separate outcome verifier can redirect the policy toward the correct task.
By Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang
arXiv:2606. 09630v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain brittle in off-nominal states requiring targeted recovery.
By Haodi Hu, Chung-Ta Huang, Jing Liu, Ye Wang, Kei Suzuki, Matthew Brand, Toshiaki Koike-Akino
The paper introduces Milestone Viability Potential Policy Optimization (MVPO), a reinforcement learning algorithm designed for long‑horizon large language model agents. MVPO addresses the zero‑credit failure problem by learning from viable failure prefixes, estimating prefix potential over Union‑Find viability regions, and repairing zero‑credit groups with potential‑difference advantages. Experiments on Qwen2.5‑1.5B‑Instruct demonstrate that MVPO outperforms eight strong baselines, improving success rates on ALFWorld and WebShop with minimal overhead.
By Qi Zhou, Yuanfan Li
arXiv:2607. 01111v1 Announce Type: cross Abstract: Robot policies inevitably encounter failures when deployed in real environments.
By Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski, Jeff Schneider