arXiv Machine Learning By Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal

Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL

Read the original on arXiv Machine Learning →

arXiv:2606. 25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.