arXiv:2608. 08604v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) is a powerful framework for solving complex collaborative tasks, but it relies heavily on well-defined global reward functions.
By Ni Mu, Yao Luan, Yiqin Yang, Qing-Shan Jia
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
arXiv:2602. 07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives.
By Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma
arXiv:2609.08211v1 Announce Type: new
Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing...
By Shanwen Mao, Hao Zhang, Guangtao nie, Zhiheng Li, Huimu Wang, Sulong Xu, Gu Simiu
arXiv:2509. 22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available.
By Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura, Mitsuki Sakamoto, Ryota Mitsuhashi, Eiji Uchibe
arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
arXiv:2607. 08647v1 Announce Type: cross Abstract: As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment.
By Ali Larian, Qian Lin, Chang Zong Wu, Daniel S. Brown
arXiv:2607. 29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases.
By Ruiming Liang, Yi Zhong, Yizhen Yuan, Yinan Zheng, Tianyi Tan, Tianyue Wang, Haiyun Guo, Jinqiao Wang, Xianyuan Zhan
arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.
By Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen
The paper introduces a diagnostic workflow for multi‑objective reinforcement learning (MORL) that reveals behavioral differences among policies on the Pareto front, which are not apparent from value vectors alone. It offers quantitative and visual tools to inspect these variations and demonstrates their effectiveness on both simple grid tasks and more complex continuous‑control benchmarks.
By Antonio Mone, Zuzanna Osika, Florian Felten, Pradeep K. Murukannaiah, Mark Fuge, Frans A. Oliehoek, Luciano Cavalcante Siebert
arXiv:2506. 13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback.
By Bilal Faye, Hanane Azzag, Mustapha Lebbah
arXiv:2606. 26397v1 Announce Type: cross Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal.
By Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia