arXiv Machine Learning

Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management

The paper presents a graph-based framework for large-scale railway network management, combining a hierarchical Bayesian model with a Gaussian Process on a graph kernel to infer spatially correlated maintenance environments from Swiss Federal Railways data. It introduces a topology-aware Multi-Agent Reinforcement Learning system that uses graph neural networks and Transformers to optimize network-level policies. The approach demonstrates scalability via zero-shot transfer learning, enabling agents trained on small network segments to perform effectively on unseen large networks, outperforming heuristics and standard MARL baselines while reducing training time.

Hugging Face Trending Papers
Sep 24

Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management

The paper presents a graph-based framework for large-scale railway network management that combines a hierarchical Bayesian model with a Gaussian Process on a graph kernel to model spatially correlated maintenance environments, and a topology-aware Multi-Agent Reinforcement Learning system using graph neural networks and Transformers to optimize network-level policies. It demonstrates scalability by training agents on small network segments and deploying them zero-shot on larger, unseen networks, achieving superior performance over heuristics and standard MARL baselines while reducing training time. The approach addresses the computational challenges of centralized methods and the coordination gaps of decentralized methods in complex, long-horizon infrastructure asset management.

arXiv AI
Jul 7

Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets

arXiv:2607. 05179v1 Announce Type: cross Abstract: In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private information about their objectives and performance, where regulatory constraints prohibit communication or direct information exchange between competitors to prevent explicit collusion.

By Enrique Adrian Villarrubia-Martin, David Mu\~noz-Valero, Luis Rodriguez-Benitez, Giovanni Montana, Luis Jimenez-Linares
arXiv Machine Learning
Sep 11

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.

By Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup
arXiv Machine Learning
Sep 21

Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks

arXiv:2609.21945v1 Announce Type: new Abstract: Urban transportation networks present complex optimization challenges spanning calibration of high-fidelity simulators and real-time operational contro...

By Adewumi Augustine Adepitan, Christopher J. Haruna, Oluwasegun Adegoke, Ayooluwatomiwa Ajiboye, Oluwatobi Oluwasakin
arXiv Machine Learning
Jul 8

GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning

arXiv:2601. 20753v4 Announce Type: replace Abstract: Preference-Conditioned Policy Learning (PCPL) in Multi-Objective Reinforcement Learning (MORL) approximates diverse Pareto-optimal solutions by conditioning a single policy on user-specified preferences, enabling run-time adaptation to arbitrary trade-offs without retraining.

By Zhiheng Jiang, Yunzhe Wang, Ryan Marr, Ellen Novoseller, Benjamin T. Files, Volkan Ustun
arXiv Machine Learning
Sep 23

Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing

The paper introduces a novel end‑to‑end, size‑agnostic graph reinforcement learning framework for the one‑dimensional bin packing problem (1D‑BPP). It models packing as a Markov decision process on an item‑compatibility graph, where a graph neural network actor‑critic policy learns to merge compatible partial bins. Empirical results on the BPPLIB benchmark show that the learned policy reduces the mean optimality gap of a constructive heuristic from 2.66 % to 2.31 %, performs competitively against other learned methods, and outperforms a state‑of‑the‑art learned solver on the hardest benchmark family.

By M. Asl{\i} Ayd{\i}n