arXiv AI

Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization

The paper introduces a Physics‑Informed Multi‑Agent Coordination framework that embeds calibrated BCMP queueing topologies into a decentralized multi‑agent reinforcement learning system for hospital patient flow. It formulates the problem as a Decentralized Partially Observable Markov Decision Process with coupled resource constraints, enabling departmental agents to negotiate patient routing and service scaling while exchanging localized action fingerprints to handle non‑stationarity. Empirical tests on MIMIC‑IV data show the approach reduces cumulative system delay compared to static Markovian models, heuristic dispatching, and independent multi‑agent baselines, all while preserving clinical safety constraints.

arXiv AI
Sep 17

CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning

CoRe-MARL is a cooperative multi-agent reinforcement learning framework designed for decentralized relief distribution networks. It models each local center as an agent in a Dec-POMDP, using a recurrent network to learn redistribution policies that reduce service gaps and improve the worst-served region. Experiments show that recurrent MAPPO outperforms independent PPO and heuristic baselines, maintaining competitive network-wide service while adapting to evolving supply and demand dynamics.

By Naimur Rahman Chowdhury, Shatabdi Sen Prapti, Md. Salehin Seyam, Limon Bin Hossain
arXiv Machine Learning
Jun 4

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

arXiv:2606. 04632v1 Announce Type: new Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung protection, and acid-base homeostasis.

By Teqi Hao, Yuxuan Fu, Xiaoyu Tan, Shaojie Shi, Bohao Lv, Yinghui Xu, Xihe Qiu
arXiv AI
Jul 21

Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making

arXiv:2607. 17038v1 Announce Type: new Abstract: This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sparse reward attribution, and dynamic environmental interaction, by designing and optimizing an intelligent agent workflow.

By Amez Amanj Ali, Kuo-Kun Tseng
arXiv Machine Learning
Sep 24

Optimization without Future Compromises? Decentralized Coordination via Collective and Reinforcement Learning

The paper introduces Hierarchical Reinforcement and Collective Learning (HRCL), a framework that combines multi‑agent reinforcement learning (MARL) with decentralized coordination. HRCL uses MARL at a high level to generate strategic guidance that limits the decision space for low‑level agents, enabling efficient short‑term coordination while considering long‑term effects. Experiments on synthetic, energy‑management, and drone‑swarm scenarios demonstrate faster convergence and significant reductions in system‑wide and individual costs compared to standalone MARL.

By Chuhao Qin, Evangelos Pournaras