arXiv AI By Celso de Melo, Zishan Feng, James Hale, Kazunori Terada, Giorgio Coricelli, Jonathan Gratch

Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models

Read the original on arXiv AI →

The study investigates how reputation, strategy, and emotional signals influence cooperation in generative AI models using the iterated prisoner's dilemma. Non‑reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT‑4o) showed cooperation shaped by all three factors, while reasoning models (Claude 4.6, Gemini 3, GPT‑5.2) relied more on strategy and reputation, displayed reduced emotional influence, and exhibited varied end‑game behaviors. These results highlight the growing sophistication and heterogeneity of AI social behavior, suggesting the need for standardized cooperation benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

The study examines how Theory of Mind (ToM) reasoning and prosocial beliefs influence large language models (LLMs) in the ultimatum game. By initializing LLM agents with Greedy, Fair, or Selfless beliefs and applying chain‑of‑thought or varying levels of ToM reasoning, the authors ran 2,700 simulations across several models, including o3‑mini and DeepSeek‑R1 Distilled Qwen 32B. Results show that ToM‑enhanced LLMs align more closely with human decision patterns, exhibit greater consistency, and achieve better negotiation outcomes, with Llama 3.3 70B producing the most belief‑consistent reasoning. whyItMatters":"The findings clarify the importance of incorporating Theory of Mind into LLMs to improve their alignment with human norms in cooperative decision‑making tasks."

By Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim
arXiv AI
Sep 21

Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

The paper examines how conversational AI, specifically ChatGPT, displays aspects of cooperative dialogue such as morality, politeness, and alignment compared to human-human conversations. Using over 26,000 multi‑turn dialogues and mixed‑effects modeling, the authors find that AI mimics the surface features of cooperation—like warmth and hedging—yet lacks the underlying social architecture that drives mutual adaptation. Key findings include a dissociation between AI’s moral output and human negotiation, a decline in linguistic convergence, and a reversal of typical human accommodation mechanisms when interacting with AI.

By Marina Mitiaeva, Lu Xiao
arXiv AI
Jul 7

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

arXiv:2604. 15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings.

By Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin
arXiv AI
Aug 28

Assessing mentalization in humans and large language models

The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.

By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
arXiv AI
Jun 9

Payoff scaling shapes cooperation in LLM agents across languages

arXiv:2601. 19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users.

By Trung-Kiet Huynh, Dao-Sy Duy-Minh, Thanh-Bang Cao, Phong-Hao Le, Hong-Dan Nguyen, Phu-Quy Nguyen-Lam, Minh-Luan Nguyen-Vo, Hong-Phat Pham, Phu-Hoa Pham, Thien-Kim Than, Chi-Nguyen Tran, Huy Tran, Gia-Thoai Tran-Le, Alessio Buscemi, Le Hong Trang, The Anh Han