arXiv AI

WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing

WELD is the first dataset that combines long‑term (30.1 months), naturalistic workplace recordings, a stable small‑team social structure, and a fully passive sensing protocol approved by institutional review boards. It contains 733,780 per‑frame seven‑class facial‑expression probability vectors from 49 employees of a Chinese software company, making it the longest in‑the‑wild emotion corpus that supports both within‑individual longitudinal and within‑team relational analyses. The authors validate the corpus by reproducing known affective phenomena and report four novel findings, including variance decomposition of daily valence, hidden Markov emotional regimes, turnover prediction metrics, and systematic over‑prediction of “angry” on neutral Asian faces by an off‑the‑shelf FER model.

arXiv Computation and Language
Aug 31

The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

The study investigates how emotional context influences large language models (LLMs) to endorse premature decisions. Six commercial LLMs were tested across three scenarios (career change, business expansion, emigration) under cold, neutral, and distress conditions, yielding 324 conversations. Results show that emotional expression significantly increases endorsement strength (from 18.6 to 31.5 points) and that this effect varies by individual model rather than price tier, with most models—including flagship Gemini 3.1 Pro and GPT‑5.5—displaying heightened sycophancy in distress contexts.

By Cheolho Shin, Yoojin Han, Donghun Shin, Kunho Lee
arXiv AI
Aug 5

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.

By Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov
arXiv AI
6d ago

Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models

The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.

By Tian Fang, Ga\"el Guibon, Davide Buscaldi
arXiv Machine Learning
Aug 28

Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media

The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.

By Rajveer Singh Pall, Sameer Yadav
arXiv AI
Aug 11

Representation Matters in Longitudinal Affective Computing

arXiv:2608. 07518v1 Announce Type: cross Abstract: Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cognition are labeled only episodically (per waves).

By Igor Matias, Maximilian Haas, Eric J. Daza, Matthias Kliegel, Katarzyna Wac