arXiv AI By Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum

People use fast and flat simulation to reason about new games

Read the original on arXiv AI →

arXiv:2510. 11503v2 Announce Type: replace-cross Abstract: Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

Assessing mentalization in humans and large language models

The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.

By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
arXiv AI
Sep 17

Clueing up LLMs with Tool-Augmented Deductive Reasoning

The paper introduces a text-based, multi-agent version of the board game Clue to test multi-step deductive reasoning in large language models (LLMs). Six LLM-based agents (GPT‑4o‑mini and Gemini‑2.5‑Flash) play turn‑based games, and a tool‑augmented approach uses a structured possibility matrix to convert implicit game state into explicit remaining possibilities, thereby offloading memory and deductive constraints from the agents. The study compares this tool‑augmented method against a baseline to assess its impact on reasoning quality and task success in a strategic reasoning environment.

By Rebecca Ansell, Autumn Toney-Wails
arXiv AI
Sep 17

What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

The paper presents a systematic mapping of recent chess research involving humans, engines, neural and reinforcement‑learning systems, large language models (LLMs), and hybrid approaches. It identifies 84 core study families and classifies them by agent type, strategic‑reasoning stages, and evaluation dimensions, highlighting a strong focus on situation assessment, evaluation, and action selection while noting gaps in planning, explanation, metacognition, and human–AI collaboration. The study also distinguishes hybrid systems by integration timing and cautions that improved human performance in evaluations does not automatically imply human–AI synergy.

By Paolo Ciancarini, Remo Pareschi