arXiv:2506. 23033v5 Announce Type: replace Abstract: Fairness audits are a key component of responsible machine-learning deployment.
By Yash Vardhan Tomar
The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.
By Vishnu Bindu Balachandran
arXiv:2610.01005v1 Announce Type: new
Abstract: As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairnes...
By Jie Tang, Chuanlong Xie, Lixing Zhu
arXiv:2601.03087v2 Announce Type: replace
Abstract: Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM...
By David Hartmann, Lena Pohlmann, Lelia Hanslik, Noah Gie{\ss}ing, Bettina Berendt, Pieter Delobelle
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
CleanScore is a black‑box audit framework that uses only scored outputs to assess whether models have been exposed to benchmark questions. It creates a public form and two fresh, independently written forms for each question, reports an interval for the public‑form advantage, and employs a private negative‑control bank with a transport radius to separate exposure from normal form mismatch. In a registered audit of five open models on 200 GSM8K and 200 ARC‑Challenge items, CleanScore found no exposure‑consistent advantage, bounding surface‑form inflation below five points, while also demonstrating how leaked items can inflate accuracy on unseen paraphrases and how planted advantages can be partially detected even after rewriting.
whyItMatters":"The study shows that CleanScore can detect and quantify exposure effects in benchmark models, providing a more nuanced understanding of model performance beyond simple accuracy scores."
By Jeffery Opoku, David Banahene