The paper introduces a scope‑conditioned generation framework that incorporates structured stereotype characteristics into prompts for large language models, aiming to improve the quality of counterspeech against online hate speech. The authors validate the method on a new, human‑curated dataset in English, Italian, and Spanish, showing significant gains over generic baselines in factuality, specificity, cogency, and effectiveness for both explicit and implicit stereotypes.
By Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata
Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generatio...
The paper investigates how to fairly compare language models across languages, noting that current evaluation methods vary widely and lack empirical validation. By training controlled monolingual models on parallel data and testing multilingual LLMs, the authors find that many normalized metrics suffer from biases due to tokenization, encoding, and orthographic differences. Instead, they recommend using sentence‑level negative log‑likelihood over semantically equivalent sequences for more reliable cross‑lingual comparisons.
By Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett
arXiv:2609.08515v1 Announce Type: cross
Abstract: As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical pr...
By Yuemei Xu, Kexin Xu, Jian Zhou, Haoyu Lu, Yequan Wang, Aishan Liu
arXiv:2607. 27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes.
By Farane Jalali Farahani, Corina Dima, Mojtaba Nayyeri, Raphael H. Heiberger, Steffen Staab
arXiv:2509. 08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation.
By Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng
The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.
By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva
arXiv:2605.28190v2 Announce Type: replace
Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embeddin...
By Manuel Frank, Haithem Afli
The paper investigates gender bias in machine translation evaluation metrics using an occupation-balanced subset of GAMBIT+ across seven English‑source language pairs, including a new German extension. It finds that masculine translations tend to receive higher scores and that biases align with stereotypical gender representations, though the strength varies by evaluator and language. The study highlights that assessing bias requires multiple dimensions beyond a single aggregate measure.
By Orfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed, Giuseppe Attanasio, Chrysoula Zerva
arXiv:2608.30902v1 Announce Type: new
Abstract: Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation imprac...
By Alessio Galatolo, Meriem Beloucif
arXiv:2607. 19243v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages.
By Alexander Manev
PolERo presents a new dataset of 3,574 Romanian question‑answer pairs from presidential transcripts, annotated for political evasion using a two‑level taxonomy of response clarity and fine‑grained evasion strategies. The study evaluates various classification methods—including TF‑IDF baselines, fine‑tuned encoders, a sliding‑window encoder, and zero/few‑shot LLM prompting—under matched conditions. Cross‑lingual transfer experiments via joint bilingual training and machine‑translation augmentation reveal that fine‑tuned encoders perform competitively, transfer is asymmetric, and ambivalent evasion categories with pragmatic cues remain the most challenging across all models.
By Gabriel Stefan, Sergiu Nisioi