arXiv AI By Neemesh Yadav, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim

DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories

Read the original on arXiv AI →

DialToM is a Theory of Mind benchmark created from naturalistic human-human dialogues, using a multiple-choice format. It introduces a State-Driven Diagnostic Probe that requires models to predict dialogue trajectories based solely on isolated mental-state profiles, without dialogue context. The evaluation shows that large language models are good at inferring mental states (Literal ToM) but struggle to use them for social forecasting (Functional ToM), while a domain expert scores 100% accuracy, highlighting a clear human‑AI gap.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 15

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

arXiv:2609.15972v1 Announce Type: cross Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...

By Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
arXiv AI
Jul 14

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

arXiv:2607. 11363v1 Announce Type: cross Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in ways that generalize to naturalistic settings.

By Roberta Rocca, Sami Boukortt, Geoff Keeling, Winnie Street
arXiv Computation and Language
Sep 2

CoMMET: A Psychologically Grounded Benchmark for Evaluating Theory of Mind in Multimodal LLMs

CoMMET is a new multimodal benchmark designed to evaluate Theory of Mind (ToM) in Multimodal Large Language Models (MLLMs). It expands beyond existing text-only, belief-focused tests by covering a wider range of mental states, incorporating moral evaluation, and enabling multi-turn, open-ended interactions. The dataset is grounded in psychological theory and provides a comprehensive assessment across different model families and sizes, revealing strengths, limitations, and future improvement directions.

By Ruirui Chen, Weifeng Jiang, Chengwei Qin, Kaiwen Wei, Yanzhen Yue, Cheston Tan