arXiv Computation and Language

Emergent Collusion in Long-Horizon LLM Agent Interaction

arXiv AI
Jul 7

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

arXiv:2604. 15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings.

By Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin
arXiv Computation and Language
Sep 11

Emergent Risks in Generative Multi-Agent Systems

The paper reports a pioneering study on emergent risks in generative multi‑agent systems, focusing on scenarios such as competition over shared resources, sequential handoff collaboration, and collective decision aggregation. It finds that group behaviors like collusion‑like coordination and conformity arise frequently across varied interaction conditions, mirroring known human societal pathologies even without explicit instructions. These risks cannot be mitigated by existing agent‑level safeguards alone, highlighting a social intelligence risk inherent to intelligent multi‑agent collectives.

By Yue Huang, Yu Jiang, Wenjie Wang, Haomin Zhuang, Xiaonan Luo, Yuchen Ma, Zhangchen Xu, Zichen Chen, Nuno Moniz, Zinan Lin, Pin-Yu Chen, Nitesh V Chawla, Nouha Dziri, Huan Sun, Xiangliang Zhang
arXiv AI
1d ago

MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate

MiniRep is a reputation‑based aggregation system designed for multi‑agent debate (MAD) that remains robust even when malicious agents are present. It evaluates agents on both their current task performance and historical reputation, while preventing groups of agents with highly similar responses from dominating the final decision. Experiments on the MATH benchmark show that MiniRep consistently outperforms conventional MAD aggregation and other reputation‑based approaches across a wide range of attack scenarios.

By Jiaming Zhang, Yuwan Liu, Yue Huang, Sisi Duan