arXiv Machine Learning By Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan, Julian F. P. Kooij

COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection

Read the original on arXiv Machine Learning →

arXiv:2606. 04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu