arXiv:2606. 04806v1 Announce Type: cross Abstract: LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior.
By Sichao Li, Sai Ma, Daniel Kilov, Secil Yanik Guyot, Zhuang Li, Seth Lazar
arXiv:2604. 11741v2 Announce Type: replace Abstract: Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under multiplayer game settings with imperfect and deceptive information.
By Keyang Zhong, Junlin Xie, Hefeng Wu, Haofeng Li, Guanbin Li
The paper introduces Cognitive Chain-of-Thought (CoCoT), a structured reasoning framework for vision‑language models that divides multimodal social reasoning into three cognitively inspired stages: Perception, Situation, and Norm. CoCoT improves performance across diverse tasks—multimodal intent disambiguation, theory of mind, social commonsense reasoning, and safety instruction following—by 5.9% to 4.6% on average. Fine‑tuning on CoCoT‑structured traces further boosts accuracy by 5–6% without explicit prompting, indicating that models internalize the structured reasoning pattern and that the approach enhances interpretability and social alignment in multimodal systems.
By Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim, Motahhare Eslami, Maarten Sap
arXiv:2608.23330v1 Announce Type: new
Abstract: Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human action...
By Jiapeng Li, Ping Wei, Wenjuan Han, Song-Chun Zhu, Lifeng Fan
arXiv:2601. 01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored.
By Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma, Gargi Chakraborty
TempCloze is a video cloze benchmark designed to evaluate visual temporal reasoning in Video-LLMs. The task presents a video’s beginning and ending clips and asks models to select the correct missing middle from four candidates, focusing on semantic, alignment, and progression aspects while minimizing appearance cues. Evaluation of 31 models shows that temporal alignment is the main challenge, with models performing better on semantic content and event progression but struggling to place events correctly in time.
By Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du