The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.
By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
The paper investigates whether in-context learning (ICL) in large language model agents reflects genuine recursive reasoning or simply statistical extrapolation. By testing LLM agents in a public goods game with manipulated historical feedback, the authors compare decision quality to a rational expectations equilibrium benchmark. They find that disrupting historical patterns eliminates the benefits of longer context, especially in highly interdependent settings, indicating that ICL behavior aligns more with statistical extrapolation than strategic reasoning.
By Yu Liu, Wenwen Li, Yifan Dou, Guangnan Ye
arXiv:2608. 06741v1 Announce Type: new Abstract: Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales.
By Han Wang, Philippe Beardsell, Boning Li, Aaron Sasmita, Shuai Li, Hongyuan Zha, Baoxiang Wang
The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.
By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
arXiv:2601. 19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users.
By Trung-Kiet Huynh, Dao-Sy Duy-Minh, Thanh-Bang Cao, Phong-Hao Le, Hong-Dan Nguyen, Phu-Quy Nguyen-Lam, Minh-Luan Nguyen-Vo, Hong-Phat Pham, Phu-Hoa Pham, Thien-Kim Than, Chi-Nguyen Tran, Huy Tran, Gia-Thoai Tran-Le, Alessio Buscemi, Le Hong Trang, The Anh Han
arXiv:2607. 27536v1 Announce Type: cross Abstract: Learning a strategic task changes more than what is directly taught: fine-tuning on one game can either enhance or degrade an agent's ability to reason in another.
By Joshua Caiata, Sreepriya Pulyassary, Xiang Li, Kate Larson
arXiv:2606. 07552v1 Announce Type: cross Abstract: Large language models exhibit innate behavioral tendencies when deployed as strategic agents -- notably a risk-averse "turtle" bias toward defensive play.
By Augustin Chan
arXiv:2606. 17657v1 Announce Type: new Abstract: People make decisions differently in strategic interactions.
By Zirui Cheng, Zeyu Shen, Thomas L. Griffiths, Peter Henderson
arXiv:2608. 12626v1 Announce Type: cross Abstract: Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals.
By Yi Wu, Zhimin Hu
Large language model (LLM)-based multi-agent systems (MAS) have demonstrated great potential in solving tasks with execution complexity, by distributing subtasks across cooperative agents. However, this divide-and-conquer paradigm falls short on decision-making tasks that are also prevalent in the real world.
arXiv:2608. 12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes.
By Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
The study examines how Theory of Mind (ToM) reasoning and prosocial beliefs influence large language models (LLMs) in the ultimatum game. By initializing LLM agents with Greedy, Fair, or Selfless beliefs and applying chain‑of‑thought or varying levels of ToM reasoning, the authors ran 2,700 simulations across several models, including o3‑mini and DeepSeek‑R1 Distilled Qwen 32B. Results show that ToM‑enhanced LLMs align more closely with human decision patterns, exhibit greater consistency, and achieve better negotiation outcomes, with Llama 3.3 70B producing the most belief‑consistent reasoning.
whyItMatters":"The findings clarify the importance of incorporating Theory of Mind into LLMs to improve their alignment with human norms in cooperative decision‑making tasks."
By Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim