arXiv AI By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

Read the original on arXiv AI →

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 7

ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

arXiv:2602. 21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks.

By Xiaoxuan Wang, Han Zhang, Haixin Wang, Yidan Shi, Ruoyan Li, Kaiqiao Han, Chenyi Tong, Haoran Deng, Renliang Sun, Alexander Taylor, Yanqiao Zhu, Jason Cong, Yizhou Sun, Wei Wang
arXiv Machine Learning
Jun 11

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

arXiv:2605. 03065v2 Announce Type: replace Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning.

By Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz
arXiv AI
3d ago

MA-WAM: Multi-Agent World-Action Model for Test-Time Planning

The paper introduces MA-WAM, a test‑time planning framework that uses a frozen multi‑agent flow policy to evaluate future joint actions by predicting their consequences while accounting for cross‑agent dependencies. Unlike naive extensions of single‑agent world models, MA‑WAM captures the interactions among simultaneous actions, enabling efficient candidate scoring. Experiments on 30 MARL benchmarks (MAMuJoCo, SMAC, MPE) show MA‑WAM improves performance by 22.0% over direct execution and 25.6% over uniform action selection, with only a 12.1 ms overhead on an A100 GPU.

By Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen, Guojie Wang, Hejun Wu