arXiv AI

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances explores whether a language model can embed a signal in natural language that another independent instance can detect without shared memory or coordination training. The study introduces a cooperative signalling game where a Sender describes two words, one hidden, and a Receiver must identify the target. Seven contemporary models from four architectural families were tested on 300 word pairs, revealing that most struggle to coordinate when signals must be undetectable, though one frontier model performs near-perfectly even after filtering, and that models can also use this capability for deliberate misdirection.

Hugging Face Trending Papers
Jul 22

Exposure is Optional: Learning Unlike Coordination in Language Models

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.

arXiv Machine Learning
Jun 17

Tacit Coordination of Large Language Models

arXiv:2601. 22184v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in multi-agent settings that require coordination without communication, from human-AI interaction to safety-critical scenarios.

By Ido Aharon, Emanuele La Malfa, Michael Wooldridge, Sarit Kraus
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv AI
Jul 20

Verbalizable Representations Form a Global Workspace in Language Models

arXiv:2607. 15495v1 Announce Type: cross Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning.

By Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey
arXiv Computation and Language
Sep 1

Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

The study examines how vision‑language models handle multi‑turn pragmatic interpretation in iterated reference games, where participants repeatedly identify novel referents using language. Researchers compared human performance with that of several models, manipulating context by varying its amount, order, and relevance. While humans consistently performed well, the models could use prior context but struggled to build relevant context for effective interpretation, indicating missing core skills for efficient linguistic collaboration.

By Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce
Hugging Face Trending Papers
Jun 25

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of this failure as task insensitivity: when faced with similar but distinct tasks, models might apply patterns learned during training and fail to solve the task at hand.