arXiv AI

CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making

arXiv AI
Aug 3

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.

By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy
arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
arXiv AI
Sep 7

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

CoSkill introduces a unified multi‑agent reinforcement learning framework that jointly trains a Reasoning Agent and a learnable Meta‑Skill Agent over a hierarchical skill library. By treating the meta‑skill workflow as a trainable agent and sharing a single backbone, CoSkill enables end‑to‑end co‑adaptation, allowing the Reasoning Agent to condition actions on retrieved task and step skills while the Meta‑Skill Agent refines those skills based on task performance. Experiments on ALFWorld and WebShop demonstrate that CoSkill outperforms prior skill‑based and RL baselines, achieving higher success rates and improved sample, asymptotic, and wall‑clock efficiency.

By Jinyuan Feng, Dongmin Li, Yiqun Chen, Yang Gao, Xing Chen, Huimu Wang, Zhiqiang Pu
arXiv AI
3d ago

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
arXiv AI
Aug 25

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.

By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv AI
22h ago

CollabFlow: Recursive Self-Improvement of Agent Collaboration

CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.

By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang