arXiv AI

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

arXiv:2607. 29602v1 Announce Type: cross Abstract: Reading a social situation often depends on behavior, not words alone.

arXiv Computation and Language
Sep 3

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.

By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
Hugging Face Trending Papers
Jul 21

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics.

Hugging Face Trending Papers
Jun 8

H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions

Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe conversations and retain information for downstream queries. Unlike traditional human-assistant settings, these environments are inherently multimodal, involve complex discourse phenomena such as anaphora and deixis, and contain asynchronous or conflicting information from multiple participants.