arXiv:2608. 03644v1 Announce Type: new Abstract: AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before.
By Maksymilian Wolski, Nicholas Hoernle, Johannes Forkel, Jakob Foerster
The paper introduces the "convention gap" as a metric for measuring implicit communication in cooperative AI, defined as the difference between predicted failure probability from literal messages and observed failure rates. Using the card game Hanabi, the authors analyze 101,000 play actions from human-human, AI-AI, and human-AI datasets, finding a +26.2pp gap in human pairs, a -0.7pp gap in AI pairs, and a +16.4pp gap in human-AI pairs, with the largest gaps occurring on plays with no hints. The study shows that convention compatibility, rather than raw AI-AI performance, may better predict an AI’s effectiveness with human partners.
By Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi Pari
arXiv:2602. 20804v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) is typically framed as a decentralised partially observable Markov decision process (Dec-POMDP), a setting whose hardness stems from two key challenges: partial observability and decentralised coordination.
By Kale-ab Tessera, Leonard Hinckeldey, Riccardo Zamboni, David Abel, Amos Storkey
arXiv:2606. 09826v1 Announce Type: cross Abstract: Vision-language model (VLM) agents are increasingly deployed in interactive game environments.
By Mingxian Lin, Shengju Qian, Yuqi Liu, Yi-Hua Huang, Yiyu Wang, Wei Huang, Yitang Li, Fan Zhang, Zeyu Hu, Lingting Zhu, Xin Wang, Xiaojuan Qi
The paper introduces ICRL4AHT, a large-scale benchmark for evaluating In-Context Reinforcement Learning (ICRL) in Ad-Hoc Teamwork (AHT) scenarios using Overcooked-V2. It provides a diverse teammate suite, a reproducible pipeline, and evaluates history-conditioned ICRL algorithms such as Algorithm Distillation and Decision-Pretrained Transformer. The results show that these methods often perform worse than random baselines and do not improve with longer horizons, underscoring the difficulty of strategic inference under partial observability in AHT.
By Yuheng Jing, Kai Li, Ziwen Zhang, Jiajun Zhang, Zeyao Ma, Jiaxi Yang, Lei Zhang, Zhe Wu, Jinmin He, Junliang Xing, Jian Cheng
arXiv:2605.14211v4 Announce Type: replace
Abstract: Long-horizon visuomotor tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrat...
By Benjamin Schneider, Xavier Schneider, Victor Zhong, Sun Sun