arXiv AI

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

arXiv:2607. 27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents.

arXiv AI
Sep 25

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

The paper introduces ICRL4AHT, a large-scale benchmark for evaluating In-Context Reinforcement Learning (ICRL) in Ad-Hoc Teamwork (AHT) scenarios using Overcooked-V2. It provides a diverse teammate suite, a reproducible pipeline, and evaluates history-conditioned ICRL algorithms such as Algorithm Distillation and Decision-Pretrained Transformer. The results show that these methods often perform worse than random baselines and do not improve with longer horizons, underscoring the difficulty of strategic inference under partial observability in AHT.

By Yuheng Jing, Kai Li, Ziwen Zhang, Jiajun Zhang, Zeyao Ma, Jiaxi Yang, Lei Zhang, Zhe Wu, Jinmin He, Junliang Xing, Jian Cheng
arXiv AI
Sep 23

Heterogeneous Robot Collaboration in Unstructured Environments with Grounded Generative Intelligence

arXiv:2510.26915v2 Announce Type: replace-cross Abstract: While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large...

By Zachary Ravichandran, Fernando Cladera, Ankit Prabhu, Jason Hughes, Carlos Nieto-Granda, Varun Murali, Camillo Taylor, George J. Pappas, Vijay Kumar
arXiv AI
Jun 17

Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models

arXiv:2504. 03991v2 Announce Type: replace-cross Abstract: Understanding how humans collaborate and communicate in teams is essential for improving human-agent teaming and AI-assisted decision-making.

By Siddharth Srikanth, Varun Bhatt, Boshen Zhang, Werner Hager, Charles Michael Lewis, Katia P. Sycara, Aaquib Tabrez, Stefanos Nikolaidis
arXiv AI
3d ago

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

The paper introduces Trajectory-guided Joint Policy Optimization (TJPO), a reinforcement‑learning framework that explicitly encourages diversity in the trajectories of large‑language‑model agents. By defining task‑specific trajectory descriptors, TJPO measures and optimizes diversity as a set‑level function, avoiding population‑based training. Experiments on Sokoban and ALFWorld demonstrate that TJPO yields diverse, interpretable behaviors while preserving strong task performance.

By Huaiyu Fu, Heng Cao, Hao Wang, Jian Ya, Tao Chen
arXiv Computation and Language
Sep 23

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

CONCAT is a training‑free framework that improves the efficiency of large language model (LLM) based multi‑agent systems by clustering agents according to their initial answers and selecting cluster leaders based on confidence. It uses a Theory‑of‑Mind‑inspired heuristic to predict collaboration benefits between leaders, then prunes communications to form an ad‑hoc network that reduces latency. Experiments on three LLMs and benchmarks show up to 2.02× higher accuracy/latency ratio than LLM‑Debate and a 50.1% latency reduction on Qwen2.5‑14B‑Instruct without task‑specific training.

By Ziyang Ma, Dingyi Zhang, Sichu Liang, Jiajia Chu, Pengfei Xia, Hui Zang, Deyu Zhou
arXiv AI
Aug 11

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

arXiv:2511. 02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools.

By Tim R. Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, Ece Kamar