arXiv Machine Learning

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

arXiv Machine Learning
1d ago

Accuracy Is Not Enough: A Cross-Architecture Audit of Demographic Bias in Deep Knowledge Tracing

The study audits demographic bias across four deep knowledge tracing architectures—DKT, DKVMN, SAKT, and AKT—using two large public datasets (Eedi and OULAD). It finds that bias is context‑dependent: socioeconomic bias is significant on Eedi, while gender bias appears on OULAD for most models. The most accurate model, AKT, also exhibits the greatest bias, and standard mitigation techniques such as reweighting and adversarial debiasing fail to reduce bias without sacrificing accuracy.

By Dang Quang Minh, Nguyen Dung Son, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
arXiv AI
Sep 11

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

The paper introduces a reference‑based bias detection method that audits hidden‑state representations of language models by encoding sentences as similarities to a fixed set of anchor sentences. This relative representation allows comparison across model variants, such as before and after fine‑tuning, and yields a metric called Representational Bias Shift (ΔB). ΔB correlates strongly with output‑level bias changes, can detect bias‑increasing checkpoints with high ROC AUC, and is computationally efficient, requiring only a few minutes and far less compute than traditional benchmarks.

By Marek Jeli\'nski, Jan Dubi\'nski, Maciej Chrabaszcz, Sebastian Cygert
arXiv Machine Learning
Aug 10

Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs

arXiv:2509. 16462v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities.

By Mina Arzaghi, Alireza Dehghanpour Farashah, Florian Carichon, Jean-Fran\c{c}ois Plante, Golnoosh Farnadi
arXiv AI
Jun 10

Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$\delta$}{delta} Alignment

arXiv:2606. 10632v1 Announce Type: cross Abstract: Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales.

By Junbo Ding, Xin Zang, Chenchen Pan, Donghao Song, Jiaxin Zhu, Danhuai Guo