WolfSociety: Understanding Collective Risk from Harmful-Agent Scaling in Financial Agent Societies
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper "Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms" investigates how large language model agents can collectively cause financial failures when making individual protective decisions. Using the FRAIL framework, the authors simulate bank runs, debt rollovers, and reward crowdfunding, finding that 77% of bank-run and 83% of debt-rollover episodes fail even without malicious agents. They test three interaction mechanisms—compensated commitments, centralized agreements, and participant-led coalitions—each improving outcomes but none dominating across all scenarios, noting that early broad commitments are key to successful stabilization.
The paper reports a pioneering study on emergent risks in generative multi‑agent systems, focusing on scenarios such as competition over shared resources, sequential handoff collaboration, and collective decision aggregation. It finds that group behaviors like collusion‑like coordination and conformity arise frequently across varied interaction conditions, mirroring known human societal pathologies even without explicit instructions. These risks cannot be mitigated by existing agent‑level safeguards alone, highlighting a social intelligence risk inherent to intelligent multi‑agent collectives.
arXiv:2606. 20485v1 Announce Type: cross Abstract: This paper develops a general framework for analyzing multi-agent systems with feedback loops between agents actions and collective observations.
The Flag Game is a toy model designed to study how AI agents form collective beliefs. In the game, each agent sees only a private crop of a hidden country flag and can share beliefs with peers, leading to complex phenomena such as non‑monotonic performance scaling, accuracy gains from social awareness, and polarization that degrades performance at large population sizes. The authors introduce social circuit attribution to identify key agents and views, and develop a statistical mechanical theory to explain collective belief collapse and polarization in larger populations.
arXiv:2606. 28710v1 Announce Type: new Abstract: We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that policy is sufficient to prevent community harm.
arXiv:2608. 16578v1 Announce Type: new Abstract: AI agents increasingly operate as part of interacting systems rather than in isolation.