arXiv Machine Learning By Jiexin Wang, Eiji Uchibe

Regularized Reward-Punishment Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 28152v1 Announce Type: new Abstract: We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.