Evaluating Language Model Bias with đ¤ Evaluate
Read the original on Hugging Face Blog âThe Flow has not summarised this story yet â read it at Hugging Face Blog.
The Flow has not summarised this story yet â read it at Hugging Face Blog.
The paper introduces GPTBIAS, a framework that uses powerful large language models like GPTâ4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the frameworkâs effectiveness and usability.
The paper demonstrates that large language model (LLM) evaluators, whether rewardâmodel based or prompted LLMâasâaâJudge, exhibit significant language bias in multilingual settings. Experiments with semantically identical instructionâresponse pairs across 23 languages reveal that lowerâresource languages receive higher scores, a bias that persists across eight openâweight evaluators and is not detectable by standard pairwise accuracy metrics. The authors link the bias to model uncertainty and language identity, showing it cannot be explained by content difficulty alone.
arXiv:2608.29921v1 Announce Type: cross Abstract: The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the mod...
arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.