arXiv Computation and Language

Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

The paper introduces CHIARO, a 1,000-sentence benchmark for contrastive emotion inference grounded in appraisal theory, where each scenario elicits a positive emotion in one person and a negative emotion in another. The dataset covers ten emotion classes and is human‑annotated. Evaluation shows that the best large language model achieves 67.3 macro‑F1, below human agreement, while existing emotion classifiers perform near chance. When used as a training signal alongside an existing emotion corpus, models improve on CHIARO and on six of ten external emotion benchmarks, demonstrating its value as a complementary training resource.

arXiv Computer Vision
4d ago

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.

By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv AI
4d ago

VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict

VISTA (Value-Informed Semantic Trust Arbitration) is a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. It uses a log-odds decomposition to separate emotion expectation from cue diagnosticity, allowing appraisal to change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone, VISTA achieves 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency, and a frozen-backbone probe reaches 0.600 macro CCC for appraisal readout versus 0.505 for emotion-only fine-tuning.

By Jiale Dai, Liuxian Ma, Xiaoke Niu, Wenjing Zhang, Huiying Zhao, Zhaoxiang Liu, Shiguo Lian, Guojie Song
arXiv AI
Sep 2

Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models

The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.

By Tian Fang, Ga\"el Guibon, Davide Buscaldi
arXiv AI
Aug 20

Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.

By Yousef Radwan