arXiv:2609.05806v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide...
By Amir Ben Khalifa, Fanny Bezancon, Amine Trabelsi, Bessam Abdulrazak
The paper investigates how large language models (LLMs) can exhibit stigma toward people with psychological conditions by examining their intermediate reasoning steps rather than just final answers. Using clinical expertise, the authors develop a framework to identify and rate stigmatizing language in LLM reasoning, distinguishing between overt prejudice and subtler biases. They also expand an existing mental health stigma benchmark to include more psychological conditions, finding that reasoning analysis reveals far more stigma than traditional multiple-choice evaluations and exposes flaws in the models’ logic and understanding of mental health.
By Sreehari Sankar, Aliakbar Nafar, Mona Barman, Hannah K. Heitz, Ashwin Kumar, Pouria Tohidi, Dailun Li, Danish Hussain, Russell DuBois, Hamed Hasheminia, Farshad Majzoubi
arXiv:2601. 05751v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to influence and persuade.
By Amalie Brogaard Pauli, Maria Barrett, Max M\"uller-Eberstein, Isabelle Augenstein, Ira Assent
The paper examines how large language models (LLMs) respond to different demographic cues—such as names—when users seek advice, focusing on race and gender in a U.S. context. It finds that using different cues for the same group leads to only partially overlapping changes in model responses, producing inconsistent conclusions about personalization and unstable bias metrics. The authors argue that LLMs react to linguistic signals tied to cues rather than to stable demographic categories, and they call for evaluations that use multiple cues and consider underlying mechanisms.
By Manuel Tonneau, Neil K. R. Sehgal, Niyati Malhotra, Sharif Kazemi, Victor Orozco-Olvera, Ana Mar\'ia Mu\~noz Boudet, Lakshmi Subramanian, Samuel P. Fraiberger, Sharath Chandra Guntuku, Valentin Hofmann
arXiv:2606. 10380v1 Announce Type: cross Abstract: Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.
By Grace Byun, Abigail Lott, Rebecca Lipschutz, Sean T. Minton, Elizabeth A. Stinson, Jinho D. Choi
The study demonstrates that large language models (LLMs) can produce highly reliable labels under a single experimental setup, yet their outputs vary significantly when researchers alter task designs or model choices. By evaluating seven LLMs across 12 task designs and 3,000 tweets for offensive language and hate speech, the authors found that agreement dropped from a median Fleiss' κ of 0.91 to a median Cohen's κ of 0.76 when task designs changed. This design sensitivity inflates prevalence estimates by up to 110.6 times compared to sampling variance alone, and confidence scores fail to mitigate the issue.
By Thomas Reiter, Christoph Kern, Fedor Miasnikov, Sofiia Nikolenko, Rob Chew, Stephanie Eckman, Frauke Kreuter