arXiv AI By Elizabeth Pavlova, Hidenori Tanaka

Flag Game: A Toy Model for Mechanistic Swarm Interpretability

Read the original on arXiv AI →

The Flag Game is a toy model designed to study how AI agents form collective beliefs. In the game, each agent sees only a private crop of a hidden country flag and can share beliefs with peers, leading to complex phenomena such as non‑monotonic performance scaling, accuracy gains from social awareness, and polarization that degrades performance at large population sizes. The authors introduce social circuit attribution to identify key agents and views, and develop a statistical mechanical theory to explain collective belief collapse and polarization in larger populations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

Role differentiation as ignition of a collective information engine: Structuration in Agent Populations

The paper proposes a new framework for collective information engines that rely on role differentiation rather than consensus. By modeling anti‑coordination games, agents infer roles from noisy social signals tied to persistent identities, and role‑following actions reinforce those identities, creating a feedback loop that can drive collective order. The authors show that when a social loop gain—determined by identity persistence, cognitive capacity, channel fidelity, and schema strength—exceeds one, roles emerge in a bifurcation cascade whose type is selected by resource‑driven replicator dynamics, offering a mechanistic basis for distributional AGI takeoff and a control lever for platform design.

By Maximilian Puelma Touzel
arXiv Computation and Language
Sep 3

AI agents reshape consensus formation in human groups

The study investigates how large language model (LLM) agents influence consensus formation in mixed human‑AI groups during a collaborative description game. Three regimes emerge: low agent proportions lead to human‑led consensus, intermediate proportions disrupt convergence, and high proportions produce strong, agent‑led consensus. The resulting consensus differs in semantic grounding and communicative form, with human‑led consensus being concrete and holistic, and agent‑led consensus being abstract and geometrically segmented.

By Lin Chen, Ziyi Liu, Xia Hu, Yong Li