arXiv AI

Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

arXiv:2608. 10810v1 Announce Type: cross Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions.

arXiv Computer Vision
4d ago

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.

By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv Computation and Language
Sep 1

MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

MMDS-Bench is a new diagnostic benchmark for multimodal dynamic stance classification in social media parent‑reply interactions. It contains 3,482 multimodal instances annotated with a seven‑label stance taxonomy, plus an 800‑instance subset that demands structured reasoning over parent and reply understanding and stance‑relation inference. The benchmark also tags each instance with five challenge factors—multimodal fusion, parent framing, non‑literal expression, interaction reasoning, and label‑boundary ambiguity—and evaluates 12 multimodal large language models using a reference‑grounded LLM‑judge protocol, revealing that current models still struggle with relational inference beyond separate parent and reply comprehension.

By Yuzhe Ding, Kang He, Li Zheng, Shengwu Zheng, Teng Shi, Fei Li, Chong Teng, Donghong Ji
Hugging Face Trending Papers
Aug 11

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations.