arXiv Machine Learning

FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

arXiv:2608. 03852v1 Announce Type: new Abstract: This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs.

arXiv Machine Learning
Sep 24

Optimization without Future Compromises? Decentralized Coordination via Collective and Reinforcement Learning

The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.

By Chuhao Qin, Evangelos Pournaras
arXiv AI
Sep 4

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

The paper proposes a deployment‑focused framework for deadline‑constrained network control, introducing the Effective Congestion (EC) metric family and Uniform Path Grouping (UPG) heuristic to better capture traffic urgency and balance load. It integrates these with a Multi‑Agent Deep Reinforcement Learning architecture (MADRL EC (p*)) that combines a distributed scheduler and a centralized RL router. A unified training objective merges live‑reward, pre‑collected‑reward, and policy‑imitation terms, leading to the Model‑Guided Annealed Reinforcement Learning (MGA‑RL) protocol built on DDPG, which generalizes offline‑to‑online learning for demonstration‑driven training.

By Vincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca
arXiv AI
Jun 4

Generalizable Multi-Task Learning for Wireless Networks Using Prompt Decision Transformers

arXiv:2606. 04328v1 Announce Type: cross Abstract: Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio resource management (RRM) toward artificial intelligence (AI)-driven RRM.

By Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci
arXiv AI
Aug 25

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.

By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv AI
Jun 11

Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling

arXiv:2501. 12942v2 Announce Type: replace Abstract: Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center management, where efficient resource allocation is required among users with diverse delay sensitivities.

By Zhuoran Li, Ruishuo Chen, Hai Zhong, Longbo Huang
arXiv Machine Learning
Sep 23

Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

The paper introduces Network Feasibility Geometry Reinforcement Learning (NFG‑RL), a method that enforces multi‑layer network constraints—such as interference, power‑rate coupling, flow conservation, service chains, capacity, latency, and reliability—by transporting a proto‑policy through a differentiable feasibility map. By compiling heterogeneous constraints into typed residual blocks and using a variational transport operator, NFG‑RL ensures almost‑sure feasible execution and shapes exploration and gradients to respect active constraints. Experiments on two wireless‑edge surrogate environments show that NFG‑RL boosts feasible utility by 37.5–41.5 %, cuts raw‑action violations by 48.5–60.8 %, and reduces P99 delay by 57.0–75.5 % compared to leading baselines.

By Zuyuan Zhang, Zeyu Fang, Mahdi Imani, Nathaniel D. Bastian, Tian Lan