arXiv AI

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It

arXiv Computation and Language
Aug 27

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models

The paper introduces Paraesthesia, a dynamic backdoor attack that uses emotionally styled inputs as triggers for large language models. By mapping target emotions into a valence–arousal space and rewriting a small subset of clean samples, the attack achieves over 98% success while minimally affecting clean performance. Experiments on four major LLMs show that the trigger cannot be fully explained by token-level cues and remains robust against several filtering and mitigation techniques.

By Ziyu Liu, Tao Li, Tao Yang, Tianjie Ni, Xiaolong Lan, Wengang Ma, Junjiang He
arXiv AI
Aug 20

Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.

By Yousef Radwan
arXiv Machine Learning
Aug 27

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

The study examines whether a decodable "empathy" direction can be used as a reliable causal lever in language models. Using EPITOME-derived facets of Recognition (cognitive) and Resonance (affective) across three instruction‑tuned LLMs, the authors find that while affective steering can partially raise affective scores, cognitive steering shows inconsistent or unmeasurable effects. The results highlight that decodability does not guarantee reliable control, especially for cognitive empathy, and that measurement sensitivity must be explicitly checked.

By Haoran Jisun
arXiv Computation and Language
Sep 1

Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative. "whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."

By Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
arXiv Machine Learning
Jun 16

Faithful Action-unit Causal Reasoning for Counterfactually Faithful Emotion Explanations

arXiv:2606. 15779v1 Announce Type: cross Abstract: Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction.

By Van Thong Huynh, Hong Hai Nguyen, Thuy Pham, Trong Nghia Nguyen, Soo-Hyung Kim
arXiv AI
2d ago

When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

The study investigates whether language‑grounded explanations improve trust in automated sheep pain recognition from facial expressions. By grounding a model in the Sheep Pain Facial Expression Scale (SPFES) and testing attention‑based explanations, the authors find that such explanations are largely ineffective. They then replace the appearance bypass with a concept bottleneck that reads only SPFES concept scores, which slightly reduces performance but yields demonstrably learned concepts and better recovery of minority pain states.

By Alam Noor, Miguel Guti'errez Gait'an