arXiv Machine Learning By DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Read the original on arXiv Machine Learning →

Large language models (LLMs) are increasingly used to assess social bias in text, but the passages they evaluate often contain surface noise such as typos and broken punctuation. This study applied five realistic noise conditions at varying intensities to 3,822 stereotype‑related responses and compared bias judgments on noisy versus original text. The findings show that noise disproportionately turns neutral judgments into biased ones—up to 120 times more likely—while rarely converting biased judgments into neutral ones, and that the most fragile LLM judge exhibits the greatest distortion at mild noise levels. As LLMs become more robust, the bias distortion tends toward parity rather than reversal, meaning bias measured on noisy text is systematically overestimated, especially in fairness‑critical categories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 10

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models used as judges for social bias are affected by noisy text, such as typos and broken punctuation. In experiments with 3,822 stereotype-related responses, noise more often turns neutral judgments into biased ones than the reverse, with up to a 120‑fold difference. The effect is strongest at mild realistic noise levels and leads to systematic overestimation of bias, especially in fairness‑critical categories.

arXiv Machine Learning
Sep 3

GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models

The paper introduces GPTBIAS, a framework that uses powerful large language models like GPT‑4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the framework’s effectiveness and usability.

By Jiaxu Zhao, Meng Fang, Shirui Pan, Wenpeng Yin, Mykola Pechenizkiy
arXiv Computation and Language
Sep 2

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.

By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed