Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
arXiv:2607. 18691v1 Announce Type: new Abstract: Progresses have been made on understanding emotion mechanisms of large language models (LLMs).
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
arXiv:2606. 14742v1 Announce Type: cross Abstract: Do LLMs have emotions?
Sofroniew et al. (2026) showed that emotion concepts in Claude Sonnet 4.5 are encoded as vectors whose geometry mirrors human affect psychology. This study replicates that finding using the base pretrained model google/gemma-2-27b, generating 205,200 Claude Sonnet 4.5 stories, extracting 171 emotion vectors, and recovering a similar affective circumplex with principal components explaining comparable variance. The analysis further identifies a sharp geometric seam at layers 22‑26, demonstrates that much of the geometry already exists in static token embeddings, and shows that the geometry predicts token‑level co‑activation with high correlation.
arXiv:2604. 07801v2 Announce Type: replace-cross Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language.
arXiv:2603.18007v2 Announce Type: replace-cross Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to inf...
arXiv:2601. 00181v3 Announce Type: replace-cross Abstract: We address two persistent gaps in Emotion Recognition in Conversation: which modeling choices materially affect performance, and how recognition findings connect to interpretable discourse-level patterns.
The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.
arXiv:2607. 00661v1 Announce Type: cross Abstract: Explanations for emotion classifiers are usually produced post hoc, with no guarantee that they reflect the computation behind the label.
arXiv:2609.22362v1 Announce Type: new Abstract: Debates about whether artificial systems can feel are often forced between two unsatisfactory positions: behavioral equivalence is treated as sufficien...
arXiv:2609.37497v1 Announce Type: new Abstract: Modern transformer models excel at capturing semantic relationships through sentence embeddings, yet their ability to perform pragmatic reasoning remai...
The paper introduces Paraesthesia, a dynamic backdoor attack that uses emotionally styled inputs as triggers for large language models. By mapping target emotions into a valence–arousal space and rewriting a small subset of clean samples, the attack achieves over 98% success while minimally affecting clean performance. Experiments on four major LLMs show that the trigger cannot be fully explained by token-level cues and remains robust against several filtering and mitigation techniques.
arXiv:2609.16247v1 Announce Type: new Abstract: Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may e...