arXiv AI By Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi Pari

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

Read the original on arXiv AI →

The paper introduces the "convention gap" as a metric for measuring implicit communication in cooperative AI, defined as the difference between predicted failure probability from literal messages and observed failure rates. Using the card game Hanabi, the authors analyze 101,000 play actions from human-human, AI-AI, and human-AI datasets, finding a +26.2pp gap in human pairs, a -0.7pp gap in AI pairs, and a +16.4pp gap in human-AI pairs, with the largest gaps occurring on plays with no hints. The study shows that convention compatibility, rather than raw AI-AI performance, may better predict an AI’s effectiveness with human partners.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 6

The Yokai Learning Environment: Tracking Beliefs Over Space and Time

arXiv:2508. 12480v3 Announce Type: replace Abstract: The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC), which evaluates an algorithm by measuring the performance of independently trained agents when paired.

By Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Foerster, Andreas Bulling
arXiv Computation and Language
Sep 21

Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry

The study evaluates six advanced language models on a two‑agent <log_2 N>‑Questions game, where a questioner must identify a secret Wikipedia paragraph using exactly <log_2 N> binary questions answered by an agent that only sees the target. Across 408 games, win rates decline geometrically with horizon length (p≈0.93), and per‑round failure rates remain flat, indicating that errors compound because more rounds must succeed rather than because individual rounds become harder. Adjudication reveals that losses stem from single‑agent answer errors and discrimination failures, with Claude Opus 5 lagging due to high false‑negative rates, while the top five models cluster closely; maximizing information gain requires structural partitioning, and neither reasoning‑token usage nor API cost correlates with success.

By Peter Potash
arXiv AI
Sep 25

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

PUBG Ally is an embodied, voice‑enabled AI teammate for PUBG: BATTLEGROUNDS that can perceive the game world, interpret player speech, and autonomously decide actions while keeping speech synchronized with gameplay. It combines a language‑model agent that uses a controlled interface to gather game information and a faster control layer for movement, combat, and recovery. The system was trained on nearly 39,000 real‑player sessions and evaluated through player feedback and preference comparisons, with live deployment requiring low‑latency on‑device execution and safety safeguards.

By Beomsoo Kim, Byeongju Kim, Dohyun Kim, Dongwon Kim, Eunchong Kim, Hongmin Kim, Hyeojung Im, Hyeonbin Hwang, Hyeonghwan Kim, Hyoseok Seol, Insub Im, Irene Chen, Jaeseung Jeon, Jimin Hong, Kiyoon Yoo, Minkyoung Park, Seohyeon Jung, Seungjun Chung, Sue Hyun Park, Sungwoo Kim, Youngin Cho, Yujeong Son, Kangwook Lee, Hyunseung Kim
arXiv Computation and Language
Sep 3

AI agents reshape consensus formation in human groups

The study investigates how large language model (LLM) agents influence consensus formation in mixed human‑AI groups during a collaborative description game. Three regimes emerge: low agent proportions lead to human‑led consensus, intermediate proportions disrupt convergence, and high proportions produce strong, agent‑led consensus. The resulting consensus differs in semantic grounding and communicative form, with human‑led consensus being concrete and holistic, and agent‑led consensus being abstract and geometrically segmented.

By Lin Chen, Ziyi Liu, Xia Hu, Yong Li
arXiv Machine Learning
4d ago

Amadeus: When Models of People Meet

arXiv:2609.35835v1 Announce Type: cross Abstract: With the sheer constant advancements raining down in the field of Artificial Intelligence, one particular possibility that may cross our mind is whet...

By Karl Hanna