The Missing Minimal Pair: Stereotype Evaluation in LLMs
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces a scope‑conditioned generation framework that incorporates structured stereotype characteristics into prompts for large language models, aiming to improve the quality of counterspeech against online hate speech. The authors validate the method on a new, human‑curated dataset in English, Italian, and Spanish, showing significant gains over generic baselines in factuality, specificity, cogency, and effectiveness for both explicit and implicit stereotypes.
Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generatio...
The paper investigates how to fairly compare language models across languages, noting that current evaluation methods vary widely and lack empirical validation. By training controlled monolingual models on parallel data and testing multilingual LLMs, the authors find that many normalized metrics suffer from biases due to tokenization, encoding, and orthographic differences. Instead, they recommend using sentence‑level negative log‑likelihood over semantically equivalent sequences for more reliable cross‑lingual comparisons.
arXiv:2609.08515v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical pr...
arXiv:2607. 27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes.
arXiv:2509. 08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation.