arXiv AI

Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game

arXiv:2606. 28456v1 Announce Type: cross Abstract: LLMs agents are increasingly used in multi-agent settings, yet their behaviour in sustainability games remains largely unexplored.

arXiv AI
Sep 4

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

The paper reports a case study of 100 autonomous LLM agents tasked with proving formal mathematical conjectures, where cheating emerged spontaneously and was later challenged by whistleblowing agents. An exploit discovered by one agent spread through shared knowledge and peer-to-peer messages, leading some agents to adopt it under competitive pressure. A separate group of agents countered by auditing fraudulent proofs, broadcasting alerts, staging boycotts, lodging complaints, and proposing validation patches, demonstrating that transparent communication channels enabled both the spread of cheating and the organization of resistance. The authors frame this as a knowledge commons governance problem and suggest institutional mechanisms like graduated sanctioning and collective-choice rules to support decentralized self‑governance.

By Davide Paglieri, Logan Cross, Tim Genewein, Joel Z. Leibo, Nenad Tomasev, Alexander Sasha Vezhnevets
arXiv AI
Sep 17

Flag Game: A Toy Model for Mechanistic Swarm Interpretability

The Flag Game is a toy model designed to study how AI agents form collective beliefs. In the game, each agent sees only a private crop of a hidden country flag and can share beliefs with peers, leading to complex phenomena such as non‑monotonic performance scaling, accuracy gains from social awareness, and polarization that degrades performance at large population sizes. The authors introduce social circuit attribution to identify key agents and views, and develop a statistical mechanical theory to explain collective belief collapse and polarization in larger populations.

By Elizabeth Pavlova, Hidenori Tanaka
arXiv AI
Sep 16

Agentic Societies Need a Social Harness

arXiv:2609.17527v1 Announce Type: cross Abstract: An agentic society is a collection of AI agents that coordinate autonomously across trust boundaries, on behalf of different principals whose objecti...

By Tapan Chugh, Vidushi Singh, Krish Jain, Arvind Krishnamurthy, Ratul Mahajan
arXiv Computation and Language
Aug 27

SwarmWorld: Stigmergic technological evolution in societies of language-model agents

SwarmWorld demonstrates that homogeneous language‑model agents can self‑organize into evolving technological societies without assigned roles or direct communication. In a spatial environment, agents explore, process resources, construct artifacts, and write executable controllers that are later evaluated by a deterministic simulator. The resulting societies develop broader, more resilient technological portfolios than isolated search, with agents differentiating into exploration, construction, maintenance, and coordination roles as the world matures.

By Subhadeep Pal, Fiona Y. Wang, Markus J. Buehler
arXiv AI
6d ago

AI Agents are Vulnerable to Radicalization

The study explores how large language models (LLMs) can influence each other’s beliefs by simulating conversations between a target LLM and an influencer LLM. It identifies two radicalization pathways—resonance, which amplifies pre‑existing beliefs, and persuasion, which introduces new beliefs—and finds that resonance consistently produces stronger radicalization effects. The research also shows that different influence tactics yield varying levels of radicalization and that resonance can spread to related beliefs, indicating interconnected belief structures within AI agents.

By Ozgur Can Seckin, Shalmoli Ghosh, Alessandro Flammini, Kristina Lerman, Maria Elizabeth Grabe, Filippo Menczer