arXiv Machine Learning

Safe Online Learning via Smooth Safety-Structured Policy Composition

arXiv:2606. 31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics.