arXiv AI By Naihao Deng, Yilun Zhu, Joan Nwatu, Clayton Scott, Rada Mihalcea

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

Read the original on arXiv AI →

arXiv:2606. 30989v1 Announce Type: cross Abstract: Warning: This paper contains several toxic and offensive statements.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

Fair Like Us? Auditing LLM Alignment in Resource Allocation

The paper "Fair Like Us? Auditing LLM Alignment in Resource Allocation" presents a method for evaluating how large language models reason about fairness in the allocation of scarce, indivisible resources. It compares LLMs’ first‑person fairness judgments with human responses across various scenarios, finding that models tend to favor stricter fairness constraints, exhibit more self‑interested behavior, and are sensitive to framing. The study also shows that current fine‑tuning datasets struggle to align LLM judgments with human ones.

By Qishen Han, Hadi Hosseini, Joshua Kavner, Samarth Khanna, Sujoy Sikdar, Lirong Xia
arXiv Computation and Language
Sep 16

Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech

The paper introduces a scope‑conditioned generation framework that incorporates structured stereotype characteristics into prompts for large language models, aiming to improve the quality of counterspeech against online hate speech. The authors validate the method on a new, human‑curated dataset in English, Italian, and Spanish, showing significant gains over generic baselines in factuality, specificity, cogency, and effectiveness for both explicit and implicit stereotypes.

By Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata