arXiv Computation and Language By Sreehari Sankar, Aliakbar Nafar, Mona Barman, Hannah K. Heitz, Ashwin Kumar, Pouria Tohidi, Dailun Li, Danish Hussain, Russell DuBois, Hamed Hasheminia, Farshad Majzoubi

Analyzing LLM Reasoning to Uncover Mental Health Stigma

Read the original on arXiv Computation and Language →

The paper investigates how large language models (LLMs) can exhibit stigma toward people with psychological conditions by examining their intermediate reasoning steps rather than just final answers. Using clinical expertise, the authors develop a framework to identify and rate stigmatizing language in LLM reasoning, distinguishing between overt prejudice and subtler biases. They also expand an existing mental health stigma benchmark to include more psychological conditions, finding that reasoning analysis reveals far more stigma than traditional multiple-choice evaluations and exposes flaws in the models’ logic and understanding of mental health.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 2

SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

arXiv:2609.01548v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound e...

By Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer
arXiv AI
Sep 18

Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation

The paper introduces behavioral coherence evaluation, a design‑time method that uses validation evidence from an established instrument to test relationships among outputs of large language models (LLMs). Using the Individual Level Abortion Stigma Scale, the authors prompted five LLMs to complete questionnaires for 627 personas and found that the models scored personas lower on self‑judgment but higher on worries about judgment, often reversing the reference direction for Black personas. Expert review highlighted that disclosure guidance from the models requires context about relationship safety, legal risk, and trusted support.

By Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad