Verifiable Social Reasoning for LLM Assistants
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evaluate an agent's ability to induce specific belief states in other agents by taking actions rather than using conversational persuasion, a capability we call Non-Conversational Planning ToM (NCP-ToM).
arXiv:2510.15144v4 Announce Type: replace Abstract: Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now appr...
SocialMaze is a new benchmark designed to evaluate large language models on social reasoning tasks that involve deep reasoning, dynamic interaction, and information uncertainty. It comprises six tasks drawn from social deduction games, everyday interactions, and digital communities, and includes automated checks and human validation to ensure data quality. Experiments with twelve LLMs reveal that stronger chain‑of‑thought reasoning improves performance on deeper inference tasks, while uncertainty consistently hurts results; targeted fine‑tuning on curated reasoning traces can markedly enhance structured social‑reasoning abilities.
arXiv:2609.15972v1 Announce Type: cross Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...
arXiv:2512. 20845v2 Announce Type: replace Abstract: LLMs have shown the capacity to improve their performance on reasoning tasks through reflecting on their mistakes, and acting with these reflections in mind.
arXiv:2608. 11624v1 Announce Type: cross Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions.