Emergence of cooperation: A reputation-modulated reinforcement learning
arXiv:2608. 20016v1 Announce Type: cross Abstract: Reputation is widely recognized as a key mechanism for sustaining cooperation.
arXiv:2604. 15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings.
arXiv:2608. 20016v1 Announce Type: cross Abstract: Reputation is widely recognized as a key mechanism for sustaining cooperation.
arXiv:2609.35928v1 Announce Type: cross Abstract: Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantl...
arXiv:2607. 04710v1 Announce Type: new Abstract: Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations.
arXiv:2609.24967v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the eme...
arXiv:2608. 12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes.
Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations. There, individual interests are misaligned with the common good and individual rationality leads to suboptimal group outcomes.
arXiv:2606. 07790v1 Announce Type: new Abstract: Multi-agent LLM systems increasingly rely on communication protocols for coordination, yet their robustness under adversarial and structural constraints remains poorly understood.
arXiv:2606. 30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be interpreted as a faithful proxy for human decision-making.
The paper introduces CURB, a reward‑shaping framework that penalizes the total variation distance between an agent’s action distributions under cooperation and defection histories, thereby preventing collusive equilibria in repeated games. By linking empirical Q‑learning collusion to Simple Penal Codes, the authors prove that any non‑trivial SPC can be neutralized, and demonstrate CURB’s effectiveness in both tabular and deep Q‑learning settings for Bertrand and Cournot competition.
The study investigates how reputation, strategy, and emotional signals influence cooperation in generative AI models using the iterated prisoner's dilemma. Non‑reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT‑4o) showed cooperation shaped by all three factors, while reasoning models (Claude 4.6, Gemini 3, GPT‑5.2) relied more on strategy and reputation, displayed reduced emotional influence, and exhibited varied end‑game behaviors. These results highlight the growing sophistication and heterogeneity of AI social behavior, suggesting the need for standardized cooperation benchmarks.
arXiv:2608.22152v1 Announce Type: new Abstract: Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than ac...
arXiv:2601. 19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users.