Evaluating Language Model Bias with đ€ Evaluate
Related stories
GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models
The paper introduces GPTBIAS, a framework that uses powerful large language models like GPTâ4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the frameworkâs effectiveness and usability.
Very Large Language Models and How to Evaluate Them
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
The paper demonstrates that large language model (LLM) evaluators, whether rewardâmodel based or prompted LLMâasâaâJudge, exhibit significant language bias in multilingual settings. Experiments with semantically identical instructionâresponse pairs across 23 languages reveal that lowerâresource languages receive higher scores, a bias that persists across eight openâweight evaluators and is not detectable by standard pairwise accuracy metrics. The authors link the bias to model uncertainty and language identity, showing it cannot be explained by content difficulty alone.
Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?
arXiv:2608.29921v1 Announce Type: cross Abstract: The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the mod...
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.
Challenges and Recommendations for LLM-as-a-Judge in Multilingual Settings and for Low-Resource Languages
arXiv:2607.02235v2 Announce Type: replace-cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to short...
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
EvalDetectBench is an open pipeline and benchmark designed to measure evaluation awareness in frontier large language models, enabling practitioners to test models against any Inspect-compatible evaluation. It includes a curated transcript suite from current frontier system-card evaluations and diverse deployment sources, and it assesses both how reliably models recognize they are being evaluated and how detectable individual benchmarks are. The benchmark addresses systematic bias by calibrating probes per model and harmonizing generator selection to correct for variance caused by model identity and prompt choice.
Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits
The paper examines whether removing declared language fields from deâidentified rĂ©sumĂ©s eliminates demographic leakage in large language models. By keeping language attributes identical and varying only unstructured prose across five ethnocultural groups and three cueâsalience levels, the authors find that nonâlanguage text still allows targetâgroup recovery (average 0.757, reaching 1.000 under high salience). They also show that evaluation designâsuch as allowing or forbidding tiesâdramatically affects LLMâasâaâjudge outcomes, underscoring the importance of evaluation protocol in bias audits.
LLJ Cards: Best practices for the Use of LLMs as Judges
arXiv:2609.24516v1 Announce Type: new Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...
PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
PADM'E is a method for synthesizing preferenceâaligned data to metaâevaluate languageâmodel (LM) evaluators of agentic behaviors. It reframes metaâevaluation as a preference judgment problem, generating criterionâbased data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.
EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting
EviSI is an evidenceâbased evaluation agent for lowâlatency simultaneous speechâtoâspeech translation. It combines Multidimensional Quality Metrics with interpreterâdeveloped criteria, using shared source evidence to assess four dimensionsâAnchor, Event, Logic, and Fluencyâwhile deduplicating verified errors before scoring. On EnglishâtoâChinese and ChineseâtoâEnglish data, EviSIâs rankings correlate strongly with human judgments, outperforming BLEU and COMET, and its multilingual extension maintains these correlations across five language directions.