Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2608.23095v1 Announce Type: new Abstract: Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain...
arXiv:2602. 04306v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has become crucial.
BiasGym is a cost‑effective, generalizable framework that injects specific biases into large language models via token‑based fine‑tuning while keeping the model frozen. It then uses two debiasing methods—Scope and Steer—to identify and suppress or redirect the components responsible for biased behavior. The framework enables consistent bias elicitation, precise localization of bias associations, and targeted debiasing without harming downstream performance, and it has been shown to reduce real‑world stereotypes such as labeling Italians as reckless drivers.
The study introduces a two‑dimensional framework to audit French news headlines, distinguishing salience framing—captured by four wording devices—from selection framing—captured by outlet‑level story form and high‑charge distributions. Using a 10,000‑headline supervision set annotated by LLMs and human arbitration, the authors classify 902,111 headlines from 25 outlets (2022‑2025) and find that salience and selection diverge yet correlate, that default thresholds inflate salience estimates, and that headlines mentioning Jews, the Far‑right, and Muslims exhibit the highest salience rates. The authors release their dataset, lexicons, and code, claiming it is the largest French headline framing audit to date.
arXiv:2608.29590v1 Announce Type: new Abstract: We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely...
The paper introduces GPTBIAS, a framework that uses powerful large language models like GPT‑4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the framework’s effectiveness and usability.