arXiv AI By Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

Read the original on arXiv AI →

arXiv:2606. 07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 3

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.

By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
arXiv AI
Aug 14

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.

By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
arXiv Computer Vision
Sep 3

Video2Reaction: Training Foundation Video Models to Predict Audience Reaction

Video2Reaction is a multimodal dataset that links short movie segments to the emotional reactions of viewers, gathered from social media comments. The dataset models reactions as distributions over categorical emotions, capturing the subjective and ambiguous nature of emotional perception. Experiments show that vision‑language models fine‑tuned with LoRA learn effectively from Video2Reaction and outperform specialized baselines, and that models pre‑fine‑tuned on this dataset transfer well to other emotion prediction tasks.

By Sidong Zhang, Trang Nguyen, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau