arXiv AI By Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 26397v1 Announce Type: cross Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 18

Pareto Q-Learning with Reward Machines

arXiv:2606. 19134v1 Announce Type: cross Abstract: We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward structure is specified by a set of reward machines (RMs).

By Arnaud Lequen, Cl\'ement Legrand-Lixon, L\'eo Sauli\`eres