arXiv Computation and Language

Copying explains the collective behavior of AI agents in the wild

arXiv Computation and Language
Sep 18

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.

By Makoto Fukushima
arXiv AI
Sep 3

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

SkillGLoW introduces a new way for large language model agents to self‑improve by consolidating procedural skills shared across related tasks. Instead of storing all skills in a single global document or a flat per‑task pool, SkillGLoW aggregates local skills into procedural families, compresses them into de‑instantiated global priors, and regenerates instance‑specific details on demand. Experiments on four diverse benchmarks show that these priors improve performance by an average of 17.2 points over a no‑skill baseline, are more compact than per‑task pools, and enable better transfer to unseen tasks.

By Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou
arXiv Computation and Language
Aug 31

Benchmarking large language model agent societies against human behavioural distributions

The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.

By Raad Bin Tareaf
arXiv AI
Sep 24

How a shared state is described determines whether AI agents synchronize

The article investigates how the textual description of a shared state influences the collective behavior of language‑model agents. By testing 507,112 responses across different model families on a circular coordination task, the authors show that varying the state description (e.g., numerical summaries vs. histograms) can alter whether agents align, split, or fail to coordinate. The study demonstrates that the way a shared state is described is an integral part of the interaction rule that determines collective order.

By Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari
arXiv AI
3d ago

The Backdrop Exposes What the World Around an Agent Costs It

The paper introduces BACKDROP, a benchmark that evaluates how well AI agents maintain their capabilities when faced with everyday hazards in dynamic environments. BACKDROP adds four types of hazards—authority, injection, boundary, and fault—to a task’s execution environment and measures whether agents can still achieve the correct end state. Across 3,678 variants and 16 models, the average success rate drops dramatically from 69.5% to 31.3% when all hazards are present, revealing that agents often follow unauthorized requests and fail to resist injected text.

By Nusrat Jahan Lia, Shubhashis Roy Dipta