arXiv AI
Aug 25

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

The study examines how Theory of Mind (ToM) reasoning and prosocial beliefs influence large language models (LLMs) in the ultimatum game. By initializing LLM agents with Greedy, Fair, or Selfless beliefs and applying chain‑of‑thought or varying levels of ToM reasoning, the authors ran 2,700 simulations across several models, including o3‑mini and DeepSeek‑R1 Distilled Qwen 32B. Results show that ToM‑enhanced LLMs align more closely with human decision patterns, exhibit greater consistency, and achieve better negotiation outcomes, with Llama 3.3 70B producing the most belief‑consistent reasoning. whyItMatters":"The findings clarify the importance of incorporating Theory of Mind into LLMs to improve their alignment with human norms in cooperative decision‑making tasks."

By Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim
arXiv Machine Learning
Jul 8

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations

arXiv:2607. 05863v1 Announce Type: new Abstract: Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations.

By Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Xin Chen, David Simchi-Levi
arXiv AI
Aug 28

Assessing mentalization in humans and large language models

The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.

By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang