arXiv:2607. 29559v1 Announce Type: new Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function.
By Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi
arXiv:2506. 13741v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions.
By Brahim Driss, Alex Davey, Riad Akrour
Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.
By Siyuan Li, Yifan Yu, Zhihao Zhang, Mengjing Chen, Fangzhou Zhu, Tao Zhong, Peng Liu, Jianye Hao
arXiv:2602. 08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems.
By Yanming Li, Xuelin Zhang, WenJie Lu, Ziye Tang, Maodong Wu, Haotian Luo, Tongtong Wu, Zijie Peng, Hongze Mi, Yibo Feng, Naiqiang Tan, Chao Huang, Lian Peng, Li Shen
The paper introduces CurriPO, a tree‑structured curriculum that adapts to diverse user reward models in AI alignment. By automatically building a curriculum that branches and reuses reward models, it addresses the problem of users whose reward models are hard to optimize, a group often underserved by conventional methods. Experiments on personalized continuous control demonstrate that CurriPO improves population satisfaction by 1.2–2.1× over the best baseline while cutting training time.
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
By Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim
arXiv:2606. 11284v1 Announce Type: cross Abstract: Real-world multi-agent systems, from traffic coordination to resource allocation, are often modeled as general-sum games where individual incentives conflict with collective welfare.
By Wongyu Lee, Francesco Lelli, Omran Ayoub, Massimo Tornatore
The paper introduces Preference-based Opponent Shaping (PBOS), a method that incorporates a preference parameter into an agent’s loss function to directly consider an opponent’s loss during strategy updates. By jointly learning strategy and preference parameters, PBOS aims to guide agents toward cooperative or competitive behaviors without relying on simple opponent predictions. Experiments on differentiable games demonstrate that PBOS enables agents to achieve better reward distributions across various environments.
By Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo
The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.
By Taehyung Kim, Jongeun Choi
arXiv:2511. 02304v2 Announce Type: replace-cross Abstract: We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution.
By Beyazit Yalcinkaya, Marcell Vazquez-Chanlatte, Ameesh Shah, Hanna Krasowski, Sanjit A. Seshia
arXiv:2507. 23604v2 Announce Type: replace Abstract: Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity.
By Tommaso Marzi, Cesare Alippi, Andrea Cini
arXiv:2606. 04284v1 Announce Type: cross Abstract: Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values.
By Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Ji-Ung Lee, Soyoung Oh, Isabel Valera, Vera Demberg