arXiv AI By Beiwen Zhang, Yongheng Liang, Guowei Zou, Haitao Wang, Liu Cong, Hejun Wu

PACT: Phenotype-Aware Contrastive Team Representation for Multi-Phenotype Grouped Ad Hoc Teamwork

Read the original on arXiv AI →

arXiv:2510. 25340v2 Announce Type: replace-cross Abstract: Learning to collaborate with various unfamiliar teammates poses a great challenge in the domain of multi-agent systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork

The paper introduces ROTATE, a regret-driven open‑ended training framework that jointly improves an Ad Hoc Teamwork (AHT) agent and an adversarial teammate generator. Unlike traditional two‑stage pipelines, ROTATE alternates between enhancing the agent and generating teammates that specifically probe its collaboration weaknesses. Experiments on Overcooked and Level‑Based Foraging show that ROTATE outperforms existing baselines on unseen teammates, setting a new benchmark for robust, generalizable teamwork.

By Caroline Wang, Arrasy Rahman, Benjamin Nativi, Jiaxun Cui, Yoonchang Sung, Peter Stone
arXiv AI
6d ago

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

The paper introduces ICRL4AHT, a large-scale benchmark for evaluating In-Context Reinforcement Learning (ICRL) in Ad-Hoc Teamwork (AHT) scenarios using Overcooked-V2. It provides a diverse teammate suite, a reproducible pipeline, and evaluates history-conditioned ICRL algorithms such as Algorithm Distillation and Decision-Pretrained Transformer. The results show that these methods often perform worse than random baselines and do not improve with longer horizons, underscoring the difficulty of strategic inference under partial observability in AHT.

By Yuheng Jing, Kai Li, Ziwen Zhang, Jiajun Zhang, Zeyao Ma, Jiaxi Yang, Lei Zhang, Zhe Wu, Jinmin He, Junliang Xing, Jian Cheng
arXiv Computation and Language
Sep 23

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

CONCAT is a training‑free framework that improves the efficiency of large language model (LLM) based multi‑agent systems by clustering agents according to their initial answers and selecting cluster leaders based on confidence. It uses a Theory‑of‑Mind‑inspired heuristic to predict collaboration benefits between leaders, then prunes communications to form an ad‑hoc network that reduces latency. Experiments on three LLMs and benchmarks show up to 2.02× higher accuracy/latency ratio than LLM‑Debate and a 50.1% latency reduction on Qwen2.5‑14B‑Instruct without task‑specific training.

By Ziyang Ma, Dingyi Zhang, Sichu Liang, Jiajia Chu, Pengfei Xia, Hui Zang, Deyu Zhou
arXiv AI
Aug 28

SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation

SIGMA is a hierarchical framework for cooperative multi‑agent reinforcement learning that addresses structured noise effects—local correlations in noise-induced decision impacts among agents with strong task dependencies. It groups agents into adaptive local structures using density‑based clustering, aggregates intra‑group representations to smooth deviations, and then applies inter‑group attention to integrate information while respecting heterogeneous contributions. Experiments on noisy‑observation StarCraft II tasks confirm that SIGMA improves robustness to observation noise without sacrificing performance in clean environments.

By Li Mingqian