arXiv AI

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

The paper introduces the "convention gap" as a metric for measuring implicit communication in cooperative AI, defined as the difference between predicted failure probability from literal messages and observed failure rates. Using the card game Hanabi, the authors analyze 101,000 play actions from human-human, AI-AI, and human-AI datasets, finding a +26.2pp gap in human pairs, a -0.7pp gap in AI pairs, and a +16.4pp gap in human-AI pairs, with the largest gaps occurring on plays with no hints. The study shows that convention compatibility, rather than raw AI-AI performance, may better predict an AI’s effectiveness with human partners.

arXiv AI
Aug 6

The Yokai Learning Environment: Tracking Beliefs Over Space and Time

arXiv:2508. 12480v3 Announce Type: replace Abstract: The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC), which evaluates an algorithm by measuring the performance of independently trained agents when paired.

By Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Foerster, Andreas Bulling
arXiv Computation and Language
Sep 21

Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry

The study evaluates six advanced language models on a two‑agent <log_2 N>‑Questions game, where a questioner must identify a secret Wikipedia paragraph using exactly <log_2 N> binary questions answered by an agent that only sees the target. Across 408 games, win rates decline geometrically with horizon length (p≈0.93), and per‑round failure rates remain flat, indicating that errors compound because more rounds must succeed rather than because individual rounds become harder. Adjudication reveals that losses stem from single‑agent answer errors and discrimination failures, with Claude Opus 5 lagging due to high false‑negative rates, while the top five models cluster closely; maximizing information gain requires structural partitioning, and neither reasoning‑token usage nor API cost correlates with success.

By Peter Potash
arXiv AI
Sep 25

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

PUBG Ally is an embodied, voice‑enabled AI teammate for PUBG: BATTLEGROUNDS that can perceive the game world, interpret player speech, and autonomously decide actions while keeping speech synchronized with gameplay. It combines a language‑model agent that uses a controlled interface to gather game information and a faster control layer for movement, combat, and recovery. The system was trained on nearly 39,000 real‑player sessions and evaluated through player feedback and preference comparisons, with live deployment requiring low‑latency on‑device execution and safety safeguards.

By Beomsoo Kim, Byeongju Kim, Dohyun Kim, Dongwon Kim, Eunchong Kim, Hongmin Kim, Hyeojung Im, Hyeonbin Hwang, Hyeonghwan Kim, Hyoseok Seol, Insub Im, Irene Chen, Jaeseung Jeon, Jimin Hong, Kiyoon Yoo, Minkyoung Park, Seohyeon Jung, Seungjun Chung, Sue Hyun Park, Sungwoo Kim, Youngin Cho, Yujeong Son, Kangwook Lee, Hyunseung Kim
arXiv Computation and Language
Sep 3

AI agents reshape consensus formation in human groups

The study investigates how large language model (LLM) agents influence consensus formation in mixed human‑AI groups during a collaborative description game. Three regimes emerge: low agent proportions lead to human‑led consensus, intermediate proportions disrupt convergence, and high proportions produce strong, agent‑led consensus. The resulting consensus differs in semantic grounding and communicative form, with human‑led consensus being concrete and holistic, and agent‑led consensus being abstract and geometrically segmented.

By Lin Chen, Ziyi Liu, Xia Hu, Yong Li
arXiv Machine Learning
4d ago

Amadeus: When Models of People Meet

arXiv:2609.35835v1 Announce Type: cross Abstract: With the sheer constant advancements raining down in the field of Artificial Intelligence, one particular possibility that may cross our mind is whet...

By Karl Hanna
arXiv AI
3d ago

Referential Uncertainty in Human--AI Collaboration

The study investigates how humans and AI collaborate on a puzzle task, focusing on referential uncertainty—when a description could refer to multiple objects. It finds that eliciting a belief distribution over candidate pieces yields better calibration and discrimination than raw action probabilities, and that precise descriptions or well‑targeted hedges significantly reduce the acceptance of wrong placements. However, the AI rarely externalizes uncertainty, and poorly targeted hedges can be counterproductive.

By Christian Poelitz, Finale Doshi-Velez, Si\^an Lindley
arXiv AI
Aug 26

Confident at the moment of action: belief miscalibration in LLM play under hidden information

The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.

By Bhushan Kashinath Joshi
arXiv Computation and Language
Sep 17

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

The study evaluates six frontier language models on a two‑agent <log(N)>‑Questions game using Wikipedia lead paragraphs. In each game a questioner must identify a target paragraph with exactly <log2 N> yes/no questions, while an answerer only sees the target and the question and replies with a single word. Across 408 games, the models perform similarly, with Claude Opus 5 winning 28 of 68 games and the top five models showing only marginal differences; win rates decline sharply with larger document sets, following a reliability parameter of 0.928 per question. "whyItMatters":"The results reveal how well language models can communicate under information asymmetry, highlighting that even top models struggle to extract a full bit per question and that reasoning token usage does not strongly predict success."

By Peter Potash
arXiv AI
Aug 11

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

arXiv:2511. 02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools.

By Tim R. Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, Ece Kamar