arXiv Computation and Language By Thomas Reiter, Christoph Kern, Fedor Miasnikov, Sofiia Nikolenko, Rob Chew, Stephanie Eckman, Frauke Kreuter

Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation

Read the original on arXiv Computation and Language →

The study demonstrates that large language models (LLMs) can produce highly reliable labels under a single experimental setup, yet their outputs vary significantly when researchers alter task designs or model choices. By evaluating seven LLMs across 12 task designs and 3,000 tweets for offensive language and hate speech, the authors found that agreement dropped from a median Fleiss' κ of 0.91 to a median Cohen's κ of 0.76 when task designs changed. This design sensitivity inflates prevalence estimates by up to 110.6 times compared to sampling variance alone, and confidence scores fail to mitigate the issue.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 27

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

The paper presents an instruction‑tuned large language model (LLM) based on Qwen3 that is fine‑tuned for hate speech mitigation by unifying 36 English hate speech datasets. The authors show that this generalist LLM achieves state‑of‑the‑art performance on in‑domain benchmarks and delivers significant gains in cross‑domain and cross‑lingual generalization, outperforming specialist encoder‑based classifiers.

By Lukas Edman, Daryna Dementieva, Alexander Fraser
arXiv Computation and Language
Sep 1

When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech

The paper introduces WSF-ARG+, a new dataset that pairs hate speech with check‑worthiness annotations, and presents an LLM‑in‑the‑loop framework to streamline the annotation process. Experiments with 12 open‑weight large language models demonstrate that the framework cuts human effort while maintaining annotation quality. The study also shows that incorporating check‑worthiness labels improves hate‑speech detection performance, boosting macro‑F1 scores for large models by up to 0.213 and averaging 0.154 across models.

By Nicol\'as Benjam\'in Ocampo, Tommaso Caselli, Davide Ceolin
arXiv Computation and Language
Sep 2

SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

arXiv:2609.01548v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound e...

By Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer