arXiv:2601. 21864v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment.
By Jinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri, Bowen Wei, Shloka Yada, Ziwei Zhu
BiasGym is a cost‑effective, generalizable framework that injects specific biases into large language models via token‑based fine‑tuning while keeping the model frozen. It then uses two debiasing methods—Scope and Steer—to identify and suppress or redirect the components responsible for biased behavior. The framework enables consistent bias elicitation, precise localization of bias associations, and targeted debiasing without harming downstream performance, and it has been shown to reduce real‑world stereotypes such as labeling Italians as reckless drivers.
By Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar, Haeun Yu, Arnav Arora, Isabelle Augenstein
arXiv:2608. 05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings.
By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
arXiv:2606. 01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback.
By Aitor Arronte Alvarez, Naiyi Xie Fincham
arXiv:2311.13892v4 Announce Type: replace-cross
Abstract: The social biases and unwelcome stereotypes revealed by pretrained language models are becoming obstacles to their application. Compared to n...
By Bingkang Shi, Xiaodan Zhang, Dehan Kong, Yulei Wu, Zongzhen Liu, Honglei Lyu, Longtao Huang
The paper introduces a bias depth score to differentiate between stable model preferences (Deep biases) and prompt‑dependent responses (Shallow biases) in large language models. By analyzing 4,442 opinion prompts across four models, it finds that only about a quarter of concentrated preferences persist after scenario reframing, indicating that most are shallow. The study shows Deep biases are more often inherited from pretraining and harder to remove through fine‑tuning or prompt‑based debiasing, highlighting the need to distinguish learned biases from prompt artifacts.
By An Vo, Vy Tuong Dang, Khai-Nguyen Nguyen, Emilio Villa-Cueva, Thamar Solorio, Anh Totti Nguyen, Daeyoung Kim
The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.
By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv:2606.12234v2 Announce Type: replace
Abstract: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the invol...
By Iuri Macocco, Pau Rodr\'iguez, Arno Blaas, Luca Zappella, Marco Baroni, Xavier Suau
arXiv:2511. 06148v4 Announce Type: replace-cross Abstract: As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased.
By Addison J. Wu, Ryan Liu, Xuechunzi Bai, Thomas L. Griffiths
The paper introduces GPTBIAS, a framework that uses powerful large language models like GPT‑4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the framework’s effectiveness and usability.
By Jiaxu Zhao, Meng Fang, Shirui Pan, Wenpeng Yin, Mykola Pechenizkiy
arXiv:2608. 14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications.
By Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz
The paper examines how large language models (LLMs) can be biased by irrelevant social contexts when evaluating teachers, using a large U.S. classroom transcript dataset. It shows that spurious contexts can shift model ratings by up to 1.48 points on a 7‑point scale and that standard mitigation methods like SFT and DPO are insufficient. The authors introduce Debiasing‑DPO, a method that combines contrastive reasoning‑augmented DPO with SFT, which reduces bias by 84% and improves predictive accuracy by 52% on Llama and Qwen Instruct models.
By Hyunji Nam, Dorottya Demszky