The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.
By Yousef Radwan
The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.
By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
The study investigates whether language‑grounded explanations improve trust in automated sheep pain recognition from facial expressions. By grounding a model in the Sheep Pain Facial Expression Scale (SPFES) and testing attention‑based explanations, the authors find that such explanations are largely ineffective. They then replace the appearance bypass with a concept bottleneck that reads only SPFES concept scores, which slightly reduces performance but yields demonstrably learned concepts and better recovery of minority pain states.
By Alam Noor, Miguel Guti'errez Gait'an
A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.
By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
arXiv:2609.13240v1 Announce Type: new
Abstract: The AffectiveArt Multidimensional Art Emotion Understanding task asks to jointly predict an artwork's fine-grained emotion (12 classes, 1549:1 head-to-...
By Jian Li, Fanfan Ji, Jinxiang Lai, Ying Tai, Jian Yang, Xiao-Tong Yuan, Chengjie Wang, Yabiao Wang
The paper investigates whether language‑grounded explanations improve trust in automated sheep pain detection from facial expressions. It finds that attention‑based explanations tied to the Sheep Pain Facial Expression Scale (SPFES) are largely ineffective, prompting the authors to replace the appearance bypass with a concept bottleneck that uses SPFES concept scores supervised by per‑region annotations. This new architecture slightly reduces performance but yields more semantically meaningful concepts, recovering minority pain states and revealing ear‑and‑eye severity orderings without explicit severity supervision, and highlights the pitfalls of pooled concept accuracy under clinical imbalance.
arXiv:2607. 08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution.
By Mayank Singal
arXiv:2609.36563v1 Announce Type: new
Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
By Yihao Qian, Runhao Zeng, Sicheng Zhao, Feng Liang, Hongmin Cai, Mingkui Tan
The study audits six vision‑language models (VLMs) to assess whether they consistently encode affective qualities of 3D shapes, using Kansei adjective pairs as affective axes. Across ten ShapeNet categories, models show moderate agreement (mean rank correlation 0.36) that is lower than geometric controls but higher than unrelated adjective pairs, with convergence varying widely by category and axis. The authors demonstrate how this audit informs a UI prototype that selectively exposes Kansei descriptors for generative design interfaces.
By Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel
EmoMed is a multimodal medical consultation agent that tailors its responses to users' emotional states—such as anxiety, confusion, or urgency—while preserving clinical accuracy. It processes text and medical images, detects affect indicators, and adjusts tone, structure, and detail accordingly. The system ensures factual reliability through a dual retrieval mechanism that combines web-based fact‑checking with an API‑connected, continuously updated medical knowledge base, and it has been evaluated across seven state‑of‑the‑art language models using comprehensive metrics, showing that emotionally adaptive responses outperform neutral baselines without sacrificing accuracy.
By Ivan Nasonov, Nikita Glazkov, Ivan Makovetskiy, Mikhail Mozikov, Daniil Sukhorukov, Andrey Savchenko, Ilya Makarov
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases.
arXiv:2605.26380v2 Announce Type: replace-cross
Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
By Jingru Chen, Yiming Liu, Mingtao Chen, Sijie Chen, Richeng Xuan, Liang Yang, Zhichao Hu, Fanyang Lu