arXiv Machine Learning By Dongsu Lee, Haoran Xu, Amy Zhang

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

Read the original on arXiv Machine Learning →

The paper introduces SCOUT, an offline multi-agent reinforcement learning framework that combines a generative behavioral prior with a decomposed value function for test-time action refinement. SCOUT uses optimal unified transport and Stein variational gradient descent to steer behavioral samples toward high-value regions, with the number of transport steps providing adaptive scaling. The authors prove a single-term KL bound on the joint soft-value gap under the individual-global-max principle and demonstrate that SCOUT outperforms existing methods on both discrete and continuous offline MARL benchmarks, including offline-to-online settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
5d ago

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

The paper introduces OMAF, an Online MARL framework that uses a one-step Transformer-based flow policy to generate coordinated actions efficiently. By replacing costly iterative sampling with a single-step action generation and a joint optimization scheme that couples softmax Q-value estimation with flow policy objectives, OMAF maintains expressive multimodal behavior while improving training speed. Experiments on 10 tasks from MPE and MAMuJoCo demonstrate up to 3.4× higher returns and 10.5× better sample efficiency compared to baseline methods.

By Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang
arXiv AI
Sep 28

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu