Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs).
PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.
By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.
By Po-han Li, Shenghui Chen, Sandeep Chinchali, Ufuk Topcu
Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive personalization demands a harder capability -- inferring what users care about from the multimodal traces they naturally leave behind.
CoMMET is a new multimodal benchmark designed to evaluate Theory of Mind (ToM) in Multimodal Large Language Models (MLLMs). It expands beyond existing text-only, belief-focused tests by covering a wider range of mental states, incorporating moral evaluation, and enabling multi-turn, open-ended interactions. The dataset is grounded in psychological theory and provides a comprehensive assessment across different model families and sizes, revealing strengths, limitations, and future improvement directions.
By Ruirui Chen, Weifeng Jiang, Chengwei Qin, Kaiwen Wei, Yanzhen Yue, Cheston Tan
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point.
arXiv:2609.22778v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require exper...
By Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang
arXiv:2606. 07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning.
By Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou
arXiv:2505. 05026v5 Announce Type: replace-cross Abstract: User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept.
By Jaehyun Jeon, Min Soo Kim, Jang Han Yoon, Sumin Shim, Yejin Choi, Hanbin Kim, Dae Hyun Kim, Youngjae Yu
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics.
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
arXiv:2507. 19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework.
By Sara Papi, Maike Z\"ufle, Marco Gaido, Beatrice Savoldi, Danni Liu, Ioannis Douros, Luisa Bentivogli, Jan Niehues