The paper introduces ICRL4AHT, a large-scale benchmark for evaluating In-Context Reinforcement Learning (ICRL) in Ad-Hoc Teamwork (AHT) scenarios using Overcooked-V2. It provides a diverse teammate suite, a reproducible pipeline, and evaluates history-conditioned ICRL algorithms such as Algorithm Distillation and Decision-Pretrained Transformer. The results show that these methods often perform worse than random baselines and do not improve with longer horizons, underscoring the difficulty of strategic inference under partial observability in AHT.
By Yuheng Jing, Kai Li, Ziwen Zhang, Jiajun Zhang, Zeyao Ma, Jiaxi Yang, Lei Zhang, Zhe Wu, Jinmin He, Junliang Xing, Jian Cheng
JaxAHT is a new open‑source library built with JAX that speeds up and standardizes research in Ad Hoc Teamwork (AHT). It offers a unified framework for generating teammates, training ego agents, and evaluating performance against unseen partners, delivering roughly 95× faster wall‑clock times than comparable PyTorch implementations. The library also supplies a diverse set of evaluation teammates for Level‑Based Foraging, Overcooked, and Hanabi, and is used to run a large‑scale benchmark that shows no single algorithm dominates and that agent modeling mainly helps in role‑based, diverse teammate settings.
By Caroline Wang, Rolando Fernandez, Zelal Su Mustafaoglu, Montek Kundan, Jiaxun Cui, Lingyun Xiao, Zhihan Wang, Di Yang Shi, Aditya Madhan, Johnny Liu, Arrasy Rahman, Peter Stone
arXiv:2504. 03991v2 Announce Type: replace-cross Abstract: Understanding how humans collaborate and communicate in teams is essential for improving human-agent teaming and AI-assisted decision-making.
By Siddharth Srikanth, Varun Bhatt, Boshen Zhang, Werner Hager, Charles Michael Lewis, Katia P. Sycara, Aaquib Tabrez, Stefanos Nikolaidis
arXiv:2511.07260v3 Announce Type: replace-cross
Abstract: Ad hoc teamwork (AHT) requires agents to collaborate with previously unseen teammates, which is crucial for many real-world applications. The...
By Hohei Chan, Xinzhi Zhang, Antao Xiang, Weinan Zhang, Mengchen Zhao
arXiv:2510. 25340v2 Announce Type: replace-cross Abstract: Learning to collaborate with various unfamiliar teammates poses a great challenge in the domain of multi-agent systems.
By Beiwen Zhang, Yongheng Liang, Guowei Zou, Haitao Wang, Liu Cong, Hejun Wu
CONCAT is a training‑free framework that improves the efficiency of large language model (LLM) based multi‑agent systems by clustering agents according to their initial answers and selecting cluster leaders based on confidence. It uses a Theory‑of‑Mind‑inspired heuristic to predict collaboration benefits between leaders, then prunes communications to form an ad‑hoc network that reduces latency. Experiments on three LLMs and benchmarks show up to 2.02× higher accuracy/latency ratio than LLM‑Debate and a 50.1% latency reduction on Qwen2.5‑14B‑Instruct without task‑specific training.
By Ziyang Ma, Dingyi Zhang, Sichu Liang, Jiajia Chu, Pengfei Xia, Hui Zang, Deyu Zhou
arXiv:2602. 17737v2 Announce Type: replace-cross Abstract: Mutual adaptation is a central challenge in human-AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior.
By Upasana Biswas, Durgesh Kalwar, Subbarao Kambhampati, Sarath Sreedharan
arXiv:2511. 02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools.
By Tim R. Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, Ece Kamar
arXiv:2606. 05793v1 Announce Type: cross Abstract: While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging.
By Hong Qian, Yuanhao Liu, Zihan Zhou, Zongbao Zhang, Hanjie Ge, Haotian Shi, Liang Dou, Xiangfeng Wang, Jingwen Yang, Aimin Zhou
arXiv:2508. 06336v2 Announce Type: replace-cross Abstract: We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork.
By Constantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer, Andreas Bulling
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv:2608. 06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting generalizability to state-of-the-art teaming research.
By Mateus Levi Sim\~oes Fernandes, Alberto Sardinha