arXiv:2609.15654v1 Announce Type: new
Abstract: Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimens...
By JuHeon Ha, Byounghan Lee, Yunseo Choi, Kyung-Ah Sohn
arXiv:2606. 15779v1 Announce Type: cross Abstract: Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction.
By Van Thong Huynh, Hong Hai Nguyen, Thuy Pham, Trong Nghia Nguyen, Soo-Hyung Kim
arXiv:2607. 15282v1 Announce Type: cross Abstract: Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state.
By Molood Arman
arXiv:2606. 30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated.
By Max Fomin, Elad David, Amit LeVi
The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.
By Yousef Radwan
arXiv:2609.16247v1 Announce Type: new
Abstract: Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may e...
By Valen Tagliabue, Leonard Dung, Cameron Berg