arXiv Machine Learning By Manpreet Singh, Rhythm Bhatia, Rahul Joshi

Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives

Read the original on arXiv Machine Learning →

The paper audits vision‑language models (CLIP) for gender bias using 1,500 artworks from the Metropolitan Museum of Art, focusing on zero‑shot logit differences for prompts like "masterpiece," "quality," and "influence." Unadjusted results show no significant gender effect and statistical equivalence across models, but high residual variance suggests that global zero‑shot metrics are largely noise‑dominated and may miss fine‑grained biases. The study underscores the need for multivariate confound control, equivalence testing, and provenance auditing when evaluating AI fairness in cultural heritage data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits

The paper examines whether removing declared language fields from de‑identified résumés eliminates demographic leakage in large language models. By keeping language attributes identical and varying only unstructured prose across five ethnocultural groups and three cue‑salience levels, the authors find that non‑language text still allows target‑group recovery (average 0.757, reaching 1.000 under high salience). They also show that evaluation design—such as allowing or forbidding ties—dramatically affects LLM‑as‑a‑judge outcomes, underscoring the importance of evaluation protocol in bias audits.

By Qiangju Chen, Yang Xiao
arXiv AI
Sep 10

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

The study examines how the design of audit questions influences perceived bias in large language models (LLMs). Using 40,726 requests across five models and three domains—hiring, lending, and medical triage—the authors find that demographic bias effects are not replicated when the audit is standardized. Instead, the audit’s construction, such as question phrasing and ordering, has a stronger impact on model responses than applicant demographics.

By Siddharth Vohra, Manikandan Ravikiran