arXiv AI By Xiaoyun Yin, Elmira Zahmat Doost, Shiwen Zhou, Garima Arya Yadav, Jamie C. Gorman

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

Read the original on arXiv AI →

arXiv:2510. 02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections rather than genuine mental states.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 3

A formal definition and meta-model for a machine theory of mind

arXiv:2606. 03471v1 Announce Type: new Abstract: This paper proposes, for the first time, a rigorous formal definition of the concept of Machine Theory of Mind, based on principles supported by evidence from cognitive psychology, neuroscience and artificial intelligence, and uses the above as a lens to examine state-of-the-art and current efforts in the field, driving a potential agenda for further research there able to "crack" the problem.

By Fabio Cuzzolin
arXiv AI
Aug 25

DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories

DialToM is a Theory of Mind benchmark created from naturalistic human-human dialogues, using a multiple-choice format. It introduces a State-Driven Diagnostic Probe that requires models to predict dialogue trajectories based solely on isolated mental-state profiles, without dialogue context. The evaluation shows that large language models are good at inferring mental states (Literal ToM) but struggle to use them for social forecasting (Functional ToM), while a domain expert scores 100% accuracy, highlighting a clear human‑AI gap.

By Neemesh Yadav, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim
arXiv AI
Sep 21

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.

By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum
arXiv AI
Aug 13

On Benchmarking Human-Like Intelligence in Machines

arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.

By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum