PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.
By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
arXiv:2601. 14569v2 Announce Type: replace-cross Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions.
By Leena Mathur, Bhaavanaa Thumu, Youssouf Kebe, Louis-Philippe Morency
arXiv:2609.22778v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require exper...
By Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang
CoMMET is a new multimodal benchmark designed to evaluate Theory of Mind (ToM) in Multimodal Large Language Models (MLLMs). It expands beyond existing text-only, belief-focused tests by covering a wider range of mental states, incorporating moral evaluation, and enabling multi-turn, open-ended interactions. The dataset is grounded in psychological theory and provides a comprehensive assessment across different model families and sizes, revealing strengths, limitations, and future improvement directions.
By Ruirui Chen, Weifeng Jiang, Chengwei Qin, Kaiwen Wei, Yanzhen Yue, Cheston Tan
arXiv:2608. 07512v1 Announce Type: cross Abstract: Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment.
By Dongsheng Hu, Tianyi Zhang, Chuang Liu, Yuan Zong Yong Li, Wenming Zheng, Xiu-xiu Zhan
Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive personalization demands a harder capability -- inferring what users care about from the multimodal traces they naturally leave behind.
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics.
HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.
By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
arXiv:2606. 07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning.
By Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point.
arXiv:2606. 15694v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content.
By Hangling Xie
arXiv:2603. 16859v2 Announce Type: replace Abstract: Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text.
By Tianyu Xie, Jinfa Huang, Yuexiao Ma, Rongfang Luo, Yan Yang, Wang Chen, Yuhui Zeng, Yixuan Zou, Qingchuan Ma, Zhiqiang Lu, Ruize Fang, Xiawu Zheng, Jiebo Luo, Rongrong Ji