arXiv AI By Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara, Steven McDonagh

Why Do Vision Language Models Struggle To Recognize Human Emotions?

Read the original on arXiv AI →

arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
4d ago

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.

By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv AI
Sep 25

Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

The paper introduces a steering‑vector‑based causal attribution framework to study how large vision‑language models (LVLMs) translate visual input into emotional narratives. By creating a specialized dataset, the authors uncover a functional decoupling in the LVLM’s three‑stage Adapt‑Aggregate‑Execute mechanism: visual emotional cues are first aggregated in middle layers via sentiment‑specific attention heads, then translated into narrative generation in deeper layers through emotion‑general pathways. Using these insights, they regulate emotional information routing to strengthen attention flow and amplify semantic activation, achieving significant performance gains on the MER‑UniBench and reducing emotional hallucinations through inference‑time intervention.

By Chengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei Tian
Hugging Face Trending Papers
Jul 23

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases.