The paper introduces OMAF, an Online MARL framework that uses a one-step Transformer-based flow policy to generate coordinated actions efficiently. By replacing costly iterative sampling with a single-step action generation and a joint optimization scheme that couples softmax Q-value estimation with flow policy objectives, OMAF maintains expressive multimodal behavior while improving training speed. Experiments on 10 tasks from MPE and MAMuJoCo demonstrate up to 3.4× higher returns and 10.5× better sample efficiency compared to baseline methods.
By Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang
arXiv:2609.15883v1 Announce Type: cross
Abstract: Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies...
By Jaehun Shon, Jinha Choi, Jongwook Jeon, Jongmin Lee
G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.
By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
arXiv:2606. 11087v1 Announce Type: cross Abstract: Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simulated and real robot control.
By Zhiyuan Zhou, Andy Peng, Charles Xu, Qiyang Li, Tobias Springenberg, Kevin Frans, Sergey Levine
arXiv:2602. 18291v2 Announce Type: replace Abstract: Online Multi-Agent Reinforcement Learning (MARL) is a prominent framework for efficient agent coordination.
By Zhuoran Li, Hai Zhong, Xun Wang, Qingxin Xia, Lihua Zhang, Longbo Huang
arXiv:2606. 08602v1 Announce Type: cross Abstract: We present an online reinforcement learning (RL) algorithm for fine-tuning flow-matching policies in continuous-control problems.
By Boshu Lei, Kostas Daniilidis, Antonio Loquercio