arXiv Machine Learning

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

The paper introduces SCOUT, an offline multi-agent reinforcement learning framework that combines a generative behavioral prior with a decomposed value function for test-time action refinement. SCOUT uses optimal unified transport and Stein variational gradient descent to steer behavioral samples toward high-value regions, with the number of transport steps providing adaptive scaling. The authors prove a single-term KL bound on the joint soft-value gap under the individual-global-max principle and demonstrate that SCOUT outperforms existing methods on both discrete and continuous offline MARL benchmarks, including offline-to-online settings.

arXiv AI
5d ago

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

The paper introduces OMAF, an Online MARL framework that uses a one-step Transformer-based flow policy to generate coordinated actions efficiently. By replacing costly iterative sampling with a single-step action generation and a joint optimization scheme that couples softmax Q-value estimation with flow policy objectives, OMAF maintains expressive multimodal behavior while improving training speed. Experiments on 10 tasks from MPE and MAMuJoCo demonstrate up to 3.4× higher returns and 10.5× better sample efficiency compared to baseline methods.

By Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang
arXiv AI
Sep 28

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
arXiv AI
Sep 25

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

The paper introduces ICRL4AHT, a large-scale benchmark for evaluating In-Context Reinforcement Learning (ICRL) in Ad-Hoc Teamwork (AHT) scenarios using Overcooked-V2. It provides a diverse teammate suite, a reproducible pipeline, and evaluates history-conditioned ICRL algorithms such as Algorithm Distillation and Decision-Pretrained Transformer. The results show that these methods often perform worse than random baselines and do not improve with longer horizons, underscoring the difficulty of strategic inference under partial observability in AHT.

By Yuheng Jing, Kai Li, Ziwen Zhang, Jiajun Zhang, Zeyao Ma, Jiaxi Yang, Lei Zhang, Zhe Wu, Jinmin He, Junliang Xing, Jian Cheng
arXiv AI
Jun 10

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning

arXiv:2606. 10613v1 Announce Type: cross Abstract: Diffusion-based Q-learning has emerged as a powerful paradigm for offline reinforcement learning, but its reliance on multi-step denoising makes both training and inference computationally expensive and brittle.

By Thanh Nguyen, Tri Ton, Hongbin Choe, Tung M. Luu, Chang D. Yoo