arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
arXiv:2505.11924v4 Announce Type: replace-cross
Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely th...
By Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu, Hui-Ying Shih, Pei-Yuan Wu
arXiv:2603.23114v2 Announce Type: replace
Abstract: A human's moral decision depends heavily on the context. Yet research on LLM morality has largely studied fixed scenarios. We address this gap by i...
By Adrian Sauter, Mona Schirmer
The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.
By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
BiasGym is a cost‑effective, generalizable framework that injects specific biases into large language models via token‑based fine‑tuning while keeping the model frozen. It then uses two debiasing methods—Scope and Steer—to identify and suppress or redirect the components responsible for biased behavior. The framework enables consistent bias elicitation, precise localization of bias associations, and targeted debiasing without harming downstream performance, and it has been shown to reduce real‑world stereotypes such as labeling Italians as reckless drivers.
By Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar, Haeun Yu, Arnav Arora, Isabelle Augenstein
arXiv:2607. 21558v1 Announce Type: new Abstract: Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode.
By Baihui Wang, Bernard Koch