arXiv Machine Learning By Leena Mathur, Bhaavanaa Thumu, Youssouf Kebe, Louis-Philippe Morency

Social Caption: Evaluating Social Understanding in Multimodal Models

Read the original on arXiv Machine Learning →

arXiv:2601. 14569v2 Announce Type: replace-cross Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.