Sequential Fairness Auditing with Limited Output Access
arXiv:2606. 30338v1 Announce Type: new Abstract: External evaluations are becoming increasingly central to the governance of AI systems.
arXiv:2606. 30338v1 Announce Type: new Abstract: External evaluations are becoming increasingly central to the governance of AI systems.
arXiv:2601. 16398v3 Announce Type: replace-cross Abstract: Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators.
arXiv:2506. 23033v5 Announce Type: replace Abstract: Fairness audits are a key component of responsible machine-learning deployment.
The study audits demographic bias across four deep knowledge tracing architectures—DKT, DKVMN, SAKT, and AKT—using two large public datasets (Eedi and OULAD). It finds that bias is context‑dependent: socioeconomic bias is significant on Eedi, while gender bias appears on OULAD for most models. The most accurate model, AKT, also exhibits the greatest bias, and standard mitigation techniques such as reweighting and adversarial debiasing fail to reduce bias without sacrificing accuracy.
arXiv:2609.16321v1 Announce Type: cross Abstract: Existing fairness analysis tools predominantly operate as post-training evaluation frameworks, requiring practitioners to complete the full model dev...
arXiv:2608. 00568v1 Announce Type: new Abstract: Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making.
arXiv:2606. 20461v1 Announce Type: new Abstract: Machine learning models have been shown to exhibit discriminatory outcomes or degraded performance for individuals at the intersection of multiple sensitive attributes, such as race and gender.
The paper introduces a reference‑based bias detection method that audits hidden‑state representations of language models by encoding sentences as similarities to a fixed set of anchor sentences. This relative representation allows comparison across model variants, such as before and after fine‑tuning, and yields a metric called Representational Bias Shift (ΔB). ΔB correlates strongly with output‑level bias changes, can detect bias‑increasing checkpoints with high ROC AUC, and is computationally efficient, requiring only a few minutes and far less compute than traditional benchmarks.
arXiv:2509. 16462v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities.
arXiv:2605. 28626v2 Announce Type: replace Abstract: Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to the latter.
arXiv:2606. 10632v1 Announce Type: cross Abstract: Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales.
arXiv:2606. 01719v1 Announce Type: cross Abstract: Machine learning models trained on sensitive data can inadvertently leak population-level information about their training distributions -- a threat known as distribution inference attack (DIA).