The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.
By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
arXiv:2510. 10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation.
By Enric Junque de Fortuny, Veronica Roberta Cappelli
The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.
By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
arXiv:2607. 09743v1 Announce Type: new Abstract: We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether their effects depend on model architecture.
By Pratyush Singh
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
The paper investigates whether large language models (LLMs) possess intrinsic value systems and how to quantify and align them. By projecting responses from 106 LLMs and 95,000 human survey profiles into a shared sociological space, the authors confirm that LLMs do have values, though these values form a concentrated, idealized core rather than mirroring human diversity. They introduce the Prior-Environment-Cognition (PEC) framework to mathematically define value expression and propose an adaptive Alignment Prescription that identifies minimal interventions—ranging from prompts to targeted parameter updates—to steer LLM values efficiently without harming general performance.
By Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu
arXiv:2607. 17946v1 Announce Type: cross Abstract: Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF).
By Saket Reddy, Andy Liu
The paper proposes a new method for evaluating AI accountability by analyzing the structural quality of a model’s defense for its decisions, using a four‑phase dialectical protocol based on Walton’s argumentation schemes and Govier’s criteria. Applied to nine large language models and 200 ambiguous moral-choice items, the study finds that models generally defend their reasoning well above the rubric minimum, though failures cluster on grounds and sufficiency and correlate with epistemic hedging. The protocol also reveals that models often present different argument schemes in justification than in reasoning, detects indefensible defenses, and highlights challenges in assessing retraction in AI alignment.
By Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert
arXiv:2608. 09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight.
By Yen-Shan Chen, Yu Chian Duan, Chih-En Kuo, Jian-Bin Wu, Yun-Nung Chen
arXiv:2606. 01552v1 Announce Type: new Abstract: Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision making under role-alignment value conflicts.
By Huayi Lai, Shichao Song, Simin Niu, Hanyu Wang, Jiawei Yang, Zhouxing Wang, Zhiqiang Yin, Xun Liang
arXiv:2510. 16380v2 Announce Type: replace-cross Abstract: As AI systems progress, we rely more on them to make decisions with us and for us.
By Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Rapha\"el Milli\`ere, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself.
whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."
By Rohit Patel, Alexandre Rezende, Steven McClain