The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
The paper introduces Paraesthesia, a dynamic backdoor attack that uses emotionally styled inputs as triggers for large language models. By mapping target emotions into a valence–arousal space and rewriting a small subset of clean samples, the attack achieves over 98% success while minimally affecting clean performance. Experiments on four major LLMs show that the trigger cannot be fully explained by token-level cues and remains robust against several filtering and mitigation techniques.
The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.
The study examines whether a decodable "empathy" direction can be used as a reliable causal lever in language models. Using EPITOME-derived facets of Recognition (cognitive) and Resonance (affective) across three instruction‑tuned LLMs, the authors find that while affective steering can partially raise affective scores, cognitive steering shows inconsistent or unmeasurable effects. The results highlight that decodability does not guarantee reliable control, especially for cognitive empathy, and that measurement sensitivity must be explicitly checked.
The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative. "whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."
Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without j...