arXiv:2608. 09366v1 Announce Type: new Abstract: Large-scale learning systems often face the challenge of balancing multiple, potentially competing objectives, such as fairness, accuracy, and latency.
By Corinna Cortes, Yishay Mansour, Mehryar Mohri
arXiv:2602. 07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives.
By Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma
The paper introduces a novel end‑to‑end, size‑agnostic graph reinforcement learning framework for the one‑dimensional bin packing problem (1D‑BPP). It models packing as a Markov decision process on an item‑compatibility graph, where a graph neural network actor‑critic policy learns to merge compatible partial bins. Empirical results on the BPPLIB benchmark show that the learned policy reduces the mean optimality gap of a constructive heuristic from 2.66 % to 2.31 %, performs competitively against other learned methods, and outperforms a state‑of‑the‑art learned solver on the hardest benchmark family.
By M. Asl{\i} Ayd{\i}n
arXiv:2606. 18105v1 Announce Type: cross Abstract: Network planning optimization is a fundamental problem across diverse domains, including transportation systems, communication networks, and power grids.
By Longlong Zhu, Jiashuo Yu, Zedi Chen, Yuhan Wu, Zhifan Jiang, Yuchen Xian, Yimeng Liu, Jiajie Su, Shaopeng Zhou, Xingyuan Li, Hongyan Liu, Xuan Liu, Dong Zhang, Chunming Wu, Xiang Chen
arXiv:2609.14968v1 Announce Type: new
Abstract: Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay betwe...
By Tiangang Li, Shi Ying, Xiangbo Tian
The paper introduces ControlG, a control‑theoretic framework for coordinating multi‑objective graph self‑supervised learning. It treats objective coordination as a temporal allocation problem, estimating each objective’s difficulty and antagonism, planning budgets with a Pareto‑aware log‑hypervolume planner, and scheduling updates via a PID controller. Experiments on nine datasets show that ControlG consistently outperforms state‑of‑the‑art baselines and provides an auditable schedule revealing which objectives drive learning.
By Karish Grover, Theodore Vasiloudis, Han Xie, Sixing Lu, Xiang Song, Christos Faloutsos
The paper presents a graph-based framework for large-scale railway network management, combining a hierarchical Bayesian model with a Gaussian Process on a graph kernel to infer spatially correlated maintenance environments from Swiss Federal Railways data. It introduces a topology-aware Multi-Agent Reinforcement Learning system that uses graph neural networks and Transformers to optimize network-level policies. The approach demonstrates scalability via zero-shot transfer learning, enabling agents trained on small network segments to perform effectively on unseen large networks, outperforming heuristics and standard MARL baselines while reducing training time.
By Giacomo Arcieri, Gregory Duth\'e, Christophe Muller, Konstantinos G. Papakonstantinou, Daniel Straub, Eleni Chatzi
Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.
By Siyuan Li, Yifan Yu, Zhihao Zhang, Mengjing Chen, Fangzhou Zhu, Tao Zhong, Peng Liu, Jianye Hao
arXiv:2607. 27953v1 Announce Type: new Abstract: Combinatorial optimization problems (COPs) underpin many real-world decisions, but their exponentially large search spaces make high-quality solutions costly to obtain.
By Shengda Gu, Kai Li, Xinyi Ke, Haobo Fu, Yifan Zhang, Jian Cheng
The paper presents a graph-based framework for large-scale railway network management that combines a hierarchical Bayesian model with a Gaussian Process on a graph kernel to model spatially correlated maintenance environments, and a topology-aware Multi-Agent Reinforcement Learning system using graph neural networks and Transformers to optimize network-level policies. It demonstrates scalability by training agents on small network segments and deploying them zero-shot on larger, unseen networks, achieving superior performance over heuristics and standard MARL baselines while reducing training time. The approach addresses the computational challenges of centralized methods and the coordination gaps of decentralized methods in complex, long-horizon infrastructure asset management.
arXiv:2609.08211v1 Announce Type: new
Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing...
By Shanwen Mao, Hao Zhang, Guangtao nie, Zhiheng Li, Huimu Wang, Sulong Xu, Gu Simiu
arXiv:2607. 29559v1 Announce Type: new Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function.
By Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi