arXiv Computation and Language By Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth Andr\'e

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

Read the original on arXiv Computation and Language →

The paper introduces EmoNet‑Face‑HQ, a fine‑grained emotion recognition benchmark that uses generated portraits and a 40‑category taxonomy to evaluate vision‑language models (VLMs). It finds that VLMs perform poorly when asked to generate responses but can match or surpass a fine‑tuned model (Empathic‑Insight‑Face) when their logits are read directly as binary queries. The study shows that the benchmark’s difficulty lies in the readout process rather than in perception, and that graded probability outputs yield better performance than simple yes/no questions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 20

Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.

By Yousef Radwan
arXiv Machine Learning
Sep 24

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.

By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
arXiv AI
Sep 18

When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

The study investigates whether language‑grounded explanations improve trust in automated sheep pain recognition from facial expressions. By grounding a model in the Sheep Pain Facial Expression Scale (SPFES) and testing attention‑based explanations, the authors find that such explanations are largely ineffective. They then replace the appearance bypass with a concept bottleneck that reads only SPFES concept scores, which slightly reduces performance but yields demonstrably learned concepts and better recovery of minority pain states.

By Alam Noor, Miguel Guti'errez Gait'an
arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
Hugging Face Trending Papers
Sep 17

When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

The paper investigates whether language‑grounded explanations improve trust in automated sheep pain detection from facial expressions. It finds that attention‑based explanations tied to the Sheep Pain Facial Expression Scale (SPFES) are largely ineffective, prompting the authors to replace the appearance bypass with a concept bottleneck that uses SPFES concept scores supervised by per‑region annotations. This new architecture slightly reduces performance but yields more semantically meaningful concepts, recovering minority pain states and revealing ear‑and‑eye severity orderings without explicit severity supervision, and highlights the pitfalls of pooled concept accuracy under clinical imbalance.