arXiv AI

The Yokai Learning Environment: Tracking Beliefs Over Space and Time

arXiv:2508. 12480v3 Announce Type: replace Abstract: The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC), which evaluates an algorithm by measuring the performance of independently trained agents when paired.

arXiv AI
Sep 12

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

The paper introduces the "convention gap" as a metric for measuring implicit communication in cooperative AI, defined as the difference between predicted failure probability from literal messages and observed failure rates. Using the card game Hanabi, the authors analyze 101,000 play actions from human-human, AI-AI, and human-AI datasets, finding a +26.2pp gap in human pairs, a -0.7pp gap in AI pairs, and a +16.4pp gap in human-AI pairs, with the largest gaps occurring on plays with no hints. The study shows that convention compatibility, rather than raw AI-AI performance, may better predict an AI’s effectiveness with human partners.

By Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi Pari
arXiv Machine Learning
Jun 16

Probing Dec-POMDP Reasoning in Cooperative MARL

arXiv:2602. 20804v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) is typically framed as a decentralised partially observable Markov decision process (Dec-POMDP), a setting whose hardness stems from two key challenges: partial observability and decentralised coordination.

By Kale-ab Tessera, Leonard Hinckeldey, Riccardo Zamboni, David Abel, Amos Storkey
arXiv AI
Sep 25

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

The paper introduces ICRL4AHT, a large-scale benchmark for evaluating In-Context Reinforcement Learning (ICRL) in Ad-Hoc Teamwork (AHT) scenarios using Overcooked-V2. It provides a diverse teammate suite, a reproducible pipeline, and evaluates history-conditioned ICRL algorithms such as Algorithm Distillation and Decision-Pretrained Transformer. The results show that these methods often perform worse than random baselines and do not improve with longer horizons, underscoring the difficulty of strategic inference under partial observability in AHT.

By Yuheng Jing, Kai Li, Ziwen Zhang, Jiajun Zhang, Zeyao Ma, Jiaxi Yang, Lei Zhang, Zhe Wu, Jinmin He, Junliang Xing, Jian Cheng
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv Machine Learning
Sep 4

LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games

The paper explores a runtime strategy-selection framework where a large language model (LLM) guides a pre‑trained reinforcement learning (RL) policy for non‑player characters (NPCs) in a Unity combat game without altering the underlying policy. Five NPC agents sharing a PPO policy were compared in a baseline setup and an LLM‑augmented setup, where a locally hosted Mistral 7B model assigns one of four tactical tags every five seconds based on live game state. Across 600 episodes against three scripted opponents, the LLM‑augmented agents more than doubled their win rate against a Balanced opponent, improved performance against an Evasive opponent, but struggled against an Aggressive opponent due to over‑reliance on encirclement; analysis of 2,430 strategy selections revealed limited zero‑shot differentiation with the model favoring Surround in 83.8% of cases.

By Hrithika Deepu Nair, Kayvan Karim
arXiv AI
3d ago

CollabFlow: Recursive Self-Improvement of Agent Collaboration

CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.

By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang
arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey