Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs).
arXiv:2601. 14569v2 Announce Type: replace-cross Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions.
By Leena Mathur, Bhaavanaa Thumu, Youssouf Kebe, Louis-Philippe Morency
CoMMET is a new multimodal benchmark designed to evaluate Theory of Mind (ToM) in Multimodal Large Language Models (MLLMs). It expands beyond existing text-only, belief-focused tests by covering a wider range of mental states, incorporating moral evaluation, and enabling multi-turn, open-ended interactions. The dataset is grounded in psychological theory and provides a comprehensive assessment across different model families and sizes, revealing strengths, limitations, and future improvement directions.
By Ruirui Chen, Weifeng Jiang, Chengwei Qin, Kaiwen Wei, Yanzhen Yue, Cheston Tan
arXiv:2608. 07512v1 Announce Type: cross Abstract: Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment.
By Dongsheng Hu, Tianyi Zhang, Chuang Liu, Yuan Zong Yong Li, Wenming Zheng, Xiu-xiu Zhan
HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.
By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
arXiv:2606. 07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning.
By Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou