arXiv:2608. 07512v1 Announce Type: cross Abstract: Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment.
By Dongsheng Hu, Tianyi Zhang, Chuang Liu, Yuan Zong Yong Li, Wenming Zheng, Xiu-xiu Zhan
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.
By Po-han Li, Shenghui Chen, Sandeep Chinchali, Ufuk Topcu
PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.
By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.
By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
arXiv:2609.22778v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require exper...
By Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang
Video2Reaction is a multimodal dataset that links short movie segments to the emotional reactions of viewers, gathered from social media comments. The dataset models reactions as distributions over categorical emotions, capturing the subjective and ambiguous nature of emotional perception. Experiments show that vision‑language models fine‑tuned with LoRA learn effectively from Video2Reaction and outperform specialized baselines, and that models pre‑fine‑tuned on this dataset transfer well to other emotion prediction tasks.
By Sidong Zhang, Trang Nguyen, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau