arXiv Computation and Language By Naomi Baes, Jemima Kang, Nick Haslam, Chris Groot, Alsa Wu, Luc Raszewski, Yulia Otmakhova

Automatic Evaluation of Mental Health Stigma in Online Communication

Read the original on arXiv Computation and Language →

The paper presents a new benchmark for automatically evaluating mental health stigma in online text, featuring a fine‑grained taxonomy that covers stigma mode, domain, and specific components across multiple mental health conditions. The authors annotate naturally occurring news and social media posts and test large language models and classifiers for sentiment, toxicity, and hate speech, finding that these models poorly capture stigma and often overpredict it without explicit rules. The benchmark, annotations, exemplar cases, and code are publicly released on GitHub.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 1

When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech

The paper introduces WSF-ARG+, a new dataset that pairs hate speech with check‑worthiness annotations, and presents an LLM‑in‑the‑loop framework to streamline the annotation process. Experiments with 12 open‑weight large language models demonstrate that the framework cuts human effort while maintaining annotation quality. The study also shows that incorporating check‑worthiness labels improves hate‑speech detection performance, boosting macro‑F1 scores for large models by up to 0.213 and averaging 0.154 across models.

By Nicol\'as Benjam\'in Ocampo, Tommaso Caselli, Davide Ceolin
arXiv Computation and Language
Sep 11

Analyzing LLM Reasoning to Uncover Mental Health Stigma

The paper investigates how large language models (LLMs) can exhibit stigma toward people with psychological conditions by examining their intermediate reasoning steps rather than just final answers. Using clinical expertise, the authors develop a framework to identify and rate stigmatizing language in LLM reasoning, distinguishing between overt prejudice and subtler biases. They also expand an existing mental health stigma benchmark to include more psychological conditions, finding that reasoning analysis reveals far more stigma than traditional multiple-choice evaluations and exposes flaws in the models’ logic and understanding of mental health.

By Sreehari Sankar, Aliakbar Nafar, Mona Barman, Hannah K. Heitz, Ashwin Kumar, Pouria Tohidi, Dailun Li, Danish Hussain, Russell DuBois, Hamed Hasheminia, Farshad Majzoubi
arXiv Computation and Language
Sep 2

SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

arXiv:2609.01548v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound e...

By Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer
arXiv Computation and Language
Sep 22

Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection

The paper introduces a diagnostic tool for distinguishing the use of misogynistic slurs from their mention in counter‑speech within code‑mixed Hinglish. It identifies evaluation artifacts in existing corpora, releases a 416‑item minimal‑pair contrast set that decorrelates slur presence and gendered register from labels, and proposes a pair‑consistency metric to assess model performance. Experiments show that even strong baselines struggle to consistently label counter‑speech pairs, while a large language model achieves perfect scores, indicating the benchmark measures genuine capability rather than exploitation of artifacts.

By Ashanvi Yadav, Shubham Bhardwaj
arXiv AI
Aug 28

Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes

The paper examines how four leading vision‑language models—LLaVA‑7B, Qwen‑VL, GPT‑4o mini, and Claude 3 Haiku—perform in detecting hateful content within memes. It evaluates the models under zero‑shot and few‑shot prompting, focusing not only on classification accuracy but also on the qualitative justifications they generate. The study highlights that these models often overlook contextual nuances, irony, and subtle cues essential for accurately identifying hate speech in memes.

By Muhammad Jawad Chowdhury, Adiba Hasan, Ishrak Hossain, Shahriar Ivan, Sabbir Ahmed