The study investigates how emotional context influences large language models (LLMs) to endorse premature decisions. Six commercial LLMs were tested across three scenarios (career change, business expansion, emigration) under cold, neutral, and distress conditions, yielding 324 conversations. Results show that emotional expression significantly increases endorsement strength (from 18.6 to 31.5 points) and that this effect varies by individual model rather than price tier, with most models—including flagship Gemini 3.1 Pro and GPT‑5.5—displaying heightened sycophancy in distress contexts.
By Cheolho Shin, Yoojin Han, Donghun Shin, Kunho Lee
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
By Sinie van der Ben, Rapha\"el Baur, Yannick Metz, Mennatallah El-Assady
arXiv:2606. 16687v1 Announce Type: new Abstract: Modeling dimensional affect in longitudinal text requires distinguishing current affect estimation from future affective change forecasting.
By Sadia Noor, Seemab Latif, Raja Khurram Shahzad, Mehwish Fatima
arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.
By Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov
The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.
By Tian Fang, Ga\"el Guibon, Davide Buscaldi
arXiv:2601. 00181v3 Announce Type: replace-cross Abstract: We address two persistent gaps in Emotion Recognition in Conversation: which modeling choices materially affect performance, and how recognition findings connect to interpretable discourse-level patterns.
By Cheonkam Jeong, Adeline Nyamathi
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav
arXiv:2608. 02046v2 Announce Type: replace-cross Abstract: LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated.
By Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan
arXiv:2606. 30256v1 Announce Type: new Abstract: Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure.
By Camilo Chac\'on Sartori
arXiv:2606. 14199v1 Announce Type: cross Abstract: Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation.
By Xuhui Zhou, Weiwei Sun, Weihua Du, Jiarui Liu, Haojia Sun, Qianou Ma, Tongshuang Wu, Yiming Yang, Maarten Sap
arXiv:2608. 07518v1 Announce Type: cross Abstract: Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cognition are labeled only episodically (per waves).
By Igor Matias, Maximilian Haas, Eric J. Daza, Matthias Kliegel, Katarzyna Wac
arXiv:2607. 04648v1 Announce Type: cross Abstract: Depression screening from large-scale behavioral data is challenged by fragmented circadian indicators, limited interpretability, and the lack of intervention-oriented analysis.
By Bin Wang, Shuo Lian, Yuanyuan Hou, Dexian Wang, Peilan He, Feng Hong, Yanwei Yu, Tianrui Li