arXiv Computation and Language By Hyunji Nam, Dorottya Demszky

Mitigating LLM biases toward spurious social contexts using direct preference optimization

Read the original on arXiv Computation and Language →

The paper examines how large language models (LLMs) can be biased by irrelevant social contexts when evaluating teachers, using a large U.S. classroom transcript dataset. It shows that spurious contexts can shift model ratings by up to 1.48 points on a 7‑point scale and that standard mitigation methods like SFT and DPO are insufficient. The authors introduce Debiasing‑DPO, a method that combines contrastive reasoning‑augmented DPO with SFT, which reduces bias by 84% and improves predictive accuracy by 52% on Llama and Qwen Instruct models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.