Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
Sofroniew et al. (2026) showed that emotion concepts in Claude Sonnet 4.5 are encoded as vectors whose geometry mirrors human affect psychology. This study replicates that finding using the base pretrained model google/gemma-2-27b, generating 205,200 Claude Sonnet 4.5 stories, extracting 171 emotion vectors, and recovering a similar affective circumplex with principal components explaining comparable variance. The analysis further identifies a sharp geometric seam at layers 22‑26, demonstrates that much of the geometry already exists in static token embeddings, and shows that the geometry predicts token‑level co‑activation with high correlation.
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.
The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.
arXiv:2606. 00129v1 Announce Type: cross Abstract: Large language models (LLMs) have emerged as powerful representation learners whose internal features increasingly align with human cognition.
arXiv:2606. 14742v1 Announce Type: cross Abstract: Do LLMs have emotions?
The paper investigates whether the layer that yields the highest probing accuracy in omni‑modal large language models is also the most effective for steering interventions. Across three independently developed models, the authors find that the best probing layers differ widely, whereas the most steerable layers consistently lie in a narrow mid‑to‑late range of the network. Using emotion as a testbed, they demonstrate a significant causal gap between probing and steering, and propose a two‑factor account linking readability and downstream plasticity to steering effectiveness.
arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.
arXiv:2609.36563v1 Announce Type: new Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.
arXiv:2609.22362v1 Announce Type: new Abstract: Debates about whether artificial systems can feel are often forced between two unsatisfactory positions: behavioral equivalence is treated as sufficien...
arXiv:2606. 07707v1 Announce Type: new Abstract: Decoding emotional states from neural signals has been typically framed as a discrete, single-label classification task based on emotionally stable stimuli, a formulation that oversimplifies the continuous, fluid, and co-occurring nature of human affect.
The study audits six vision‑language models (VLMs) to assess whether they consistently encode affective qualities of 3D shapes, using Kansei adjective pairs as affective axes. Across ten ShapeNet categories, models show moderate agreement (mean rank correlation 0.36) that is lower than geometric controls but higher than unrelated adjective pairs, with convergence varying widely by category and axis. The authors demonstrate how this audit informs a UI prototype that selectively exposes Kansei descriptors for generative design interfaces.